EDBT 2026 Demo / reviewers in the wild / expert
Scott Cohen
dblp:54/4155
· DBLP profile ↗
79ranked-venue papers
1as first author
21since 2021 · last 2025
0000-0002-3459-6899ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 70 · 1 first-author · 18 since 2021Artificial intelligence and machine learning · 65 · 1 first-author · 18 since 2021Databases, data management, data science and information retrieval · 5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Polarized Color Screen MattingabstractThis paper considers the long-standing problem of extracting alpha mattes from video using a known background. While various color-based or polarization-based approaches have been studied in past decades, the problem remains ill-posed because the solutions solely rely on either color or polarization. We introduce Polarized Color Screen Matting, a single-shot, per-pixel matting theory for alpha matte and foreground color recovery using both color and polarization cues. Through a theoretical analysis of our diffuse-specular polarimetric compositing equation, we derive practical closed-form matting methods with their solvability conditions. Our theory concludes that an alpha matte can be extracted without manual corrections using off-the-shelf equipment such as an LCD monitor, polarization camera, and unpolarized lights with calibrated color. Experiments on synthetic and real-world datasets verify the validity of our theory and show the capability of our matting methods on real videos with quantitative and qualitative comparisons to color-based and polarization-based matting methods. Kenji Enomoto, Scott Cohen, Brian L. Price, T. J. Rhodes |
CVPR | 2 |
| 2025 | The Photographer's Eye: Teaching Multimodal Large Language Models to See, and Critique Like Photographersabstract"While editing directly from life, photographers have found it too difficult to see simultaneously both the blue and the sky."John Szarkowski, William Eggleston’s Guide1Photographer and curator, Szarkowski insightfully revealed one of the notable gaps between general and aesthetic visual understanding: while the former focuses on identifying the factual element in an image (sky), the latter transcends such object identification, viewing it instead as an aesthetic component—a pure color block (blue). Such fundamental distinctions between general (detection, localization, etc.) and aesthetic (color, lighting, composition, etc.) visual understanding present a significant challenge for Multimodal Large Language Models (MLLMs). Although some recent works have made initial explorations, they are often limited to general and basic aesthetic commonsense. As a result, they frequently fall short in real-world scenarios (Fig. 1), which require extensive expertise—including photographic techniques, photo pre/post-processing knowledge, and more, to provide a detailed analysis and description. To fundamentally enhance the aesthetics understanding of MLLMs, we first introduce a novel dataset, PhotoCritique, derived from extensive discussions among professional photographers and enthusiasts, and characterized by the large scale, expertise, and diversity. Then, to better learn visual aesthetics from PhotoCritique, we furthur propose a novel model, PhotoEye, featuring a language-guided multi-view vision fusion mechanism to understand image aesthetics from multiple perspectives. Finally, we present a novel benchmark, PhotoBench, a comprehensive and professional benchmark for aesthetic visual understanding. On existing benchmarks and PhotoBench, our model demonstrates clear advantages over existing models. Daiqing Qi, Handong Zhao, Jing Shi 0005, Simon Jenni, Franck Dernoncourt, Scott Cohen, Sheng Li 0001 |
CVPR | 7 |
| 2025 | MetaShadow: Object-Centered Shadow Detection, Removal, and SynthesisabstractShadows are often under-considered or even ignored in image editing applications, limiting the realism of the edited results. In this paper, we introduce MetaShadow, a three-in-one versatile framework that enables detection, removal, and controllable synthesis of shadows in natural images in an object-centered fashion. MetaShadow combines the strengths of two cooperative components: Shadow Analyzer, for object-centered shadow detection and removal, and Shadow Synthesizer, for reference-based controllable shadow synthesis. Notably, we optimize the learning of the intermediate features from Shadow Analyzer to guide Shadow Synthesizer to generate more realistic shadows that blend seamlessly with the scene. Extensive evaluations on multiple shadow benchmark datasets show significant improvements of MetaShadow over the existing state-of-the-art methods on object-centered shadow detection, removal, and synthesis. MetaShadow excels in image-editing tasks such as object removal, relocation, and insertion, pushing the boundaries of object-centered image editing. Tianyu Wang 0003, Jianming Zhang 0001, Haitian Zheng, Zhihong Ding, Scott Cohen, Zhe Lin 0001, Wei Xiong 0008, Chi-Wing Fu, Luis Figueroa, Soo Ye Kim |
CVPR | 5 |
| 2025 | Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic AlignmentabstractPersonalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models. Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Zhe Lin 0001, Brian L. Price, Scott Cohen, Jianming Zhang 0001, Daniel G. Aliaga |
ICLR | 9 |
| 2024 | FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset DeduplicationabstractRecent dataset deduplication techniques have demonstrated that content-aware dataset pruning can dramatically reduce the cost of training Vision-Language Pre-trained (VLP) models without significant performance losses compared to training on the original dataset. These results have been based on pruning commonly used image-caption datasets collected from the web - datasets that are known to harbor harmful social biases that may then be codified in trained models. In this work, we evaluate how deduplication affects the prevalence of these biases in the resulting trained models and introduce an easy-to-implement modification to the recent SemDeDup algorithm that can reduce the negative effects that we observe. When examining CLIP-style models trained on deduplicated variants of LAION-400M, we find our proposed FairDeDup algorithm consistently leads to improved fairness metrics over SemDeDup on the FairFace and FACET datasets while maintaining zero-shot performance on CLIP benchmarks. Eric Slyman, Stefan Lee, Scott Cohen, Kushal Kafle |
CVPR | 3 |
| 2024 | IMPRINT: Generative Object Compositing by Learning Identity-Preserving RepresentationabstractGenerative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality. Project page: https://song630.github.io/IMPRINT-Project-Page/ Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Daniel G. Aliaga |
CVPR | 4 |
| 2024 | FineMatch: Aspect-Based Fine-Grained Image and Text Mismatch Detection and Correction
Hang Hua, Jing Shi 0005, Kushal Kafle, Simon Jenni, Daoan Zhang, John P. Collomosse, Scott Cohen, Jiebo Luo 0001 |
ECCV (9) | 7 |
| 2024 | Latent Feature-Guided Diffusion Models for Shadow RemovalabstractRecovering textures under shadows has remained a challenging problem due to the difficulty of inferring shadow-free scenes from shadow images. In this paper, we propose the use of diffusion models as they offer a promising approach to gradually refine the details of shadow regions during the diffusion process. Our method improves this process by conditioning on a learned latent feature space that inherits the characteristics of shadow-free images, thus avoiding the limitation of conventional methods that condition on degraded images only. Additionally, we propose to alleviate potential local optima during training by fusing noise features with the diffusion network. We demonstrate the effectiveness of our approach which outperforms the previous best method by 13% in terms of RMSE on the AISTD dataset. Further, we explore instance-level shadow removal, where our model outperforms the previous best method by 82% in terms of RMSE on the DESOBA dataset. Kangfu Mei, Luis Figueroa, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Vishal M. Patel |
WACV | 5 |
| 2024 | SCoRD: Subject-Conditional Relation Detection with Text-Augmented DataabstractWe propose Subject-Conditional Relation Detection (SCoRD), where conditioned on an input subject, the goal is to predict all its relations to other objects in a scene along with their locations. Based on the Open Images dataset, we propose a challenging OIv6-SCoRD benchmark such that the training and testing splits have a distribution shift in terms of the occurrence statistics of subject, relation, object triplets. To solve this problem, we propose an auto-regressive model that given a subject, it predicts its relations, objects, and object locations by casting this output as a sequence of tokens. First, we show that previous scene-graph prediction methods fail to produce as exhaustive an enumeration of relation-object pairs when conditioned on a subject on this benchmark. Particularly, we obtain a recall@3 of 83.8% for our relation-object predictions compared to the 49.75% obtained by a recent scene graph detector. Then, we show improved generalization on both relation-object and object-box predictions by leveraging during training relation-object pairs obtained automatically from textual captions and for which no object-box annotations are available. Particularly, for subject, relation, object triplets for which no object locations are available during training, we are able to obtain a recall@3 of 33.80% for relation-object pairs and 26.75% for their box locations. Kushal Kafle, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Vicente Ordonez |
WACV | 4 |
| 2024 | Structure-Guided Image Completion With Image-Level and Object-Level Semantic DiscriminatorsabstractStructure-guided image completion aims to inpaint a local region of an image according to an input guidance map from users. While such a task enables many practical applications for interactive editing, existing methods often struggle to hallucinate realistic object instances in complex natural scenes. Such a limitation is partially due to the lack of semantic-level constraints inside the hole region as well as the lack of a mechanism to enforce realistic object generation. In this work, we propose a learning paradigm that consists of semantic discriminators and object-level discriminators for improving the generation of complex semantics and objects. Specifically, the semantic discriminators leverage pretrained visual features to improve the realism of the generated visual concepts. Moreover, the object-level discriminators take aligned instances as inputs to enforce the realism of individual objects. Our proposed scheme significantly improves the generation quality and achieves state-of-the-art results on various tasks, including segmentation-guided completion, edge-guided manipulation and panoptically-guided manipulation on Places2 datasets. Furthermore, our trained model is flexible and can support multiple editing use cases, such as object insertion, replacement, removal and standard inpainting. In particular, our trained model combined with a novel automatic image completion pipeline achieves state-of-the-art results on the standard inpainting task. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Qing Liu 0017, Sohrab Amirghodsi, Yuqian Zhou, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | GamutMLP: A Lightweight MLP for Color Loss RecoveryabstractCameras and image-editing software often process images in the wide-gamut ProPhoto color space, encompassing 90% of all visible colors. However, when images are encoded for sharing, this color-rich representation is transformed and clipped to fit within the small-gamut standard RGB (sRGB) color space, representing only 30% of visible colors. Recovering the lost color information is challenging due to the clipping procedure. Inspired by neural implicit representations for 2D images, we propose a method that optimizes a lightweight multi-layer-perceptron (MLP) model during the gamut reduction step to predict the clipped values. GamutMLP takes approximately 2 seconds to optimize and requires only 23 KB of storage. The small memory footprint allows our GamutMLP model to be saved as metadata in the sRGB image—the model can be extracted when needed to restore wide-gamut color values. We demonstrate the effectiveness of our approach for color recovery and compare it with alternative strategies, including pre-trained DNN-based gamut expansion networks and other implicit neural representation methods. As part of this effort, we introduce a new color gamut dataset of 2200 wide-gamut/small-gamut images for training and testing. Hoang Minh Le 0001, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 3 |
| 2023 | ObjectStitch: Object Compositing with Diffusion ModelabstractObject compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial manual effort from professionals, and is hardly scalable. Thus, with the recent advances in generative models, in this work, we propose a selfsupervised framework for object compositing by leveraging the power of conditional diffusion models. Our framework can hollistically address the object compositing task in a unified model, transforming the viewpoint, geometry, color and shadow of the generated object while requiring no manual labeling. To preserve the input object's characteristics, we introduce a content adaptor that helps to maintain categori-cal semantics and object appearance. A data augmentation method is further adopted to improve the fidelity of the generator. Our method outperforms relevant baselines in both realism and faithfulness of the synthesized result images in a user study on various real-world images. Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, Daniel G. Aliaga |
CVPR | 4 |
| 2023 | TopNet: Transformer-Based Object Placement Network for Image CompositingabstractWe investigate the problem of automatically placing an object into a background image for image compositing. Given a background image and a segmented object, the goal is to train a model to predict plausible placements (location and scale) of the object for compositing. The quality of the composite image highly depends on the predicted location/scale. Existing works either generate candidate bounding boxes or apply sliding-window search using global representations from background and object images, which fail to model local information in background images. However, local clues in background images are important to determine the compatibility of placing the objects with certain locations/scales. In this paper, we propose to learn the correlation between object features and all local background features with a transformer module so that detailed information can be provided on all possible location/scale configurations. A sparse contrastive loss is further proposed to train our model with sparse supervision. Our new formulation generates a 3D heatmap indicating the plausibility of all location/scale combinations in one network forward pass, which is > 10 x faster than the previous sliding-window method. It also supports interactive search when users provide a pre-defined location or scale. The proposed method can be trained with explicit annotation or in a self-supervised manner using an off-the-shelf inpainting model, and it outperforms state-of-the-art methods significantly. User study shows that the trained model generalizes well to real-world images with diverse challenging scenes and object categories. Sijie Zhu, Zhe Lin 0001, Scott Cohen, Jason Kuen, Chen Chen 0001 |
CVPR | 3 |
| 2023 | Semantic Layout Manipulation With High-Resolution Sparse AttentionabstractWe tackle the problem of semantic image layout manipulation, which aims to manipulate an input image by editing its semantic label map. A core problem of this task is how to transfer visual details from the input images to the new semantic layout while making the resulting image visually realistic. Recent work on learning cross-domain correspondence has shown promising results for global layout transfer with dense attention-based warping. However, this method tends to lose texture details due to the resolution limitation and the lack of smoothness constraint on correspondence. To adapt this paradigm for the layout manipulation task, we propose a high-resolution sparse attention module that effectively transfers visual details to new layouts at a resolution up to 512x512. To further improve visual quality, we introduce a novel generator architecture consisting of a semantic encoder and a two-stage decoder for coarse-to-fine synthesis. Experiments on the ADE20k and Places365 datasets demonstrate that our proposed approach achieves substantial improvements over the existing inpainting and layout manipulation methods. Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Jianming Zhang 0001, Ning Xu 0007, Jiebo Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2022 | COAT: Correspondence-driven Object Appearance Transfer
Sangryul Jeon, Zhe Lin 0001, Scott Cohen, Zhihong Ding, Kwanghoon Sohn |
BMVC | 4 |
| 2022 | Improving Closed and Open-Vocabulary Attribute Prediction Using Transformers
Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava |
ECCV (25) | 5 |
| 2022 | Image Inpainting with Cascaded Modulation GAN and Object-Aware Training
Haitian Zheng, Zhe Lin 0001, Jingwan Lu, Scott Cohen, Eli Shechtman, Connelly Barnes, Jianming Zhang 0001, Ning Xu 0007, Sohrab Amirghodsi, Jiebo Luo 0001 |
ECCV (16) | 4 |
| 2022 | GALA: Toward Geometry-and-Lighting-Aware Object Search for Compositing
Sijie Zhu, Zhe Lin 0001, Scott Cohen, Jason Kuen, Chen Chen 0001 |
ECCV (27) | 3 |
| 2021 | Learning To Predict Visual Attributes in the WildabstractVisual attributes constitute a large portion of information contained in a scene. Objects can be described using a wide variety of attributes which portray their visual appearance (color, texture), geometry (shape, size, posture), and other intrinsic properties (state, action). Existing work is mostly limited to study of attribute prediction in specific domains. In this paper, we introduce a large-scale in-the-wild visual attribute prediction dataset consisting of over 927K attribute annotations for over 260K object instances. Formally, object attribute prediction is a multi-label classification problem where all attributes that apply to an object must be predicted. Our dataset poses significant challenges to existing methods due to large number of attributes, label sparsity, data imbalance, and object occlusion. To this end, we propose several techniques that systematically tackle these challenges, including a base model that utilizes both low- and high-level CNN features with multi-hop attention, reweighting and resampling techniques, a novel negative label expansion scheme, and a novel supervised attribute-aware contrastive learning algorithm. Using these techniques, we achieve near 3.7 mAP and 5.7 overall F1 points improvement over the current state of the art. Further details about the VAW dataset can be found at https://vawdataset.com/ Khoi Pham, Kushal Kafle, Zhe Lin 0001, Zhihong Ding, Scott Cohen, Quan Tran, Abhinav Shrivastava |
CVPR | 5 |
| 2021 | AESOP: Abstract Encoding of Stories, Objects, and Pictures
Hareesh Ravi, Kushal Kafle, Scott Cohen, Jonathan Brandt, Mubbasir Kapadia |
ICCV | 3 |
| 2021 | Deep Interactive Thin Object SelectionabstractExisting deep learning based interactive segmentation methods have achieved remarkable performance with only a few user clicks, e.g. DEXTR [32] attaining 91.5% IoU on PASCAL VOC with only four extreme clicks. However, we observe even the state-of-the-art methods would often struggle in cases of objects to be segmented with elongated thin structures (e.g. bug legs and bicycle spokes). We investigate such failures, and find the critical reasons behind are two-fold: 1) lack of appropriate training dataset; and 2) extremely imbalanced distribution w.r.t. number of pixels belonging to thin and non-thin regions. Targeted at these challenges, we collect a large-scale dataset specifically for segmentation of thin elongated objects, named ThinObject-5K. Also, we present a novel integrative thin object segmentation network consisting of three streams. Among them, the high-resolution edge stream aims at preserving fine-grained details including elongated thin parts; the fixed-resolution context stream focuses on capturing semantic contexts. The two streams' outputs are then amalgamated in the fusion stream to complement each other for help producing a refined segmentation output with sharper predictions around thin parts. Extensive experimental results well demonstrate the effectiveness of our proposed solution on segmenting thin objects, surpassing the baseline by ~ 30% IoUthindespite using only four clicks. Codes and dataset are available at https://github.com/liewjunhao/thin-object-selection. Jun Hao Liew, Scott Cohen, Brian L. Price, Long Mai, Jiashi Feng |
WACV | 2 |
| 2020 | PhraseCut: Language-Based Image Segmentation in the WildabstractWe consider the problem of segmenting image regions given a natural language phrase, and study it on a novel dataset of 77,262 images and 345,486 phrase-region pairs. Our dataset is collected on top of the Visual Genome dataset and uses the existing annotations to generate a challenging set of referring phrases for which the corresponding regions are manually annotated. Phrases in our dataset correspond to multiple regions and describe a large number of object and stuff categories as well as their attributes such as color, shape, parts, and relationships with other entities in the image. Our experiments show that the scale and diversity of concepts in our dataset poses significant challenges to the existing state-of-the-art. We systematically handle the long-tail nature of these concepts and present a modular approach to combine category, attribute, and relationship cues that outperforms existing approaches. Chenyun Wu, Zhe Lin 0001, Scott Cohen, Trung Bui, Subhransu Maji |
CVPR | 3 |
| 2020 | Deepstrip: High-Resolution Boundary RefinementabstractIn this paper, we target refining the boundaries in high resolution images given low resolution masks. For memory and computation efficiency, we propose to convert the regions of interest into strip images and compute a boundary prediction in the strip domain. To detect the target boundary, we present a framework with two prediction layers. First, all potential boundaries are predicted as an initial prediction and then a selection layer is used to pick the target boundary and smooth the result. To encourage accurate prediction, a loss which measures the boundary distance in strip domain is introduced. In addition, we enforce a matching consistency and C0 continuity regularization to the network to reduce false alarms. Extensive experiments on both public and a newly created high resolution dataset strongly validate our approach. Peng Zhou 0009, Brian L. Price, Scott Cohen, Gregg Wilensky, Larry Davis 0001 |
CVPR | 3 |
| 2020 | PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click
Henghui Ding, Scott Cohen, Brian L. Price, Xudong Jiang 0001 |
ECCV (3) | 2 |
| 2020 | Interactive Training And Architecture For Deep Object SelectionabstractInteractive object cutout tools are the cornerstone of the image editing workflow. Algorithms that can reduce the number of interactions are clearly valuable. Recent deep-learning based interactive segmentation algorithms are capable of rough binary selections with a handful of clicks, yet, they tend to plateau once this rough selection has been reached. In this work, we interpret this plateau as an inability of the algorithm to precisely leverage each user interaction.We introduce a novel interactive architecture and a training scheme that are both tailored to better exploit the user input at higher numbers of clicks. Comprehensive experiments support our approach, and our network achieves state of the art performance. Marco Forte, Brian L. Price, Scott Cohen, Ning Xu 0007, François Pitié |
ICME | 3 |
| 2020 | Figure Captioning with Relation Maps for ReasoningabstractFigures, such as line plots, pie charts, bar charts, are widely used to convey important information in a concise format. In this work, we investigate the problem of figure caption generation where the goal is to automatically generate a natural language description for a given figure. While natural image captioning has been studied extensively, figure captioning has received relatively little attention and remains a challenging problem. A successful solution to this task has many potential applications, such as: 1) automatic parsing large amount of figures in PDF document; 2) improving user experience by allowing figure content to be accessible to those with visual impairment. To solve this problem, we introduce a dataset FigCAP and propose novel attention mechanism. In order to solve the exposure bias issue, we further train the captioning model with sequence-level policy based on reinforcement learning, which directly optimizes evaluation metrics. Extensive experiments show that the proposed method outperforms the baselines, thus demonstrating a significant potential for automatic generating captions for figures. Ruiyi Zhang 0002, Eunyee Koh, Sungchul Kim, Scott Cohen, Ryan Rossi |
WACV | 5 |
| 2020 | Answering Questions about Data Visualizations using Efficient Bimodal FusionabstractChart question answering (CQA) is a newly proposed visual question answering (VQA) task where an algorithm must answer questions about data visualizations, e.g. bar charts, pie charts, and line graphs. CQA requires capabilities that natural-image VQA algorithms lack: fine-grained measurements, optical character recognition, and handling out-of-vocabulary words in both questions and answers. Without modifications, state-of-the-art VQA algorithms perform poorly on this task. Here, we propose a novel CQA algorithm called parallel recurrent fusion of image and language (PReFIL). PReFIL first learns bimodal embeddings by fusing question and image features and then intelligently aggregates these learned embeddings to answer the given question. Despite its simplicity, PReFIL greatly surpasses state-of-the art systems and human baselines on both the FigureQA and DVQA datasets. Additionally, we demonstrate that PReFIL can be used to reconstruct tables by asking a series of questions about a chart. Kushal Kafle, Robik Shrestha, Brian L. Price, Scott Cohen, Christopher Kanan |
WACV | 4 |
| 2019 | When Color Constancy Goes Wrong: Correcting Improperly White-Balanced ImagesabstractThis paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after decades of computational color constancy research, there are no effective solutions to this problem. The challenge lies not in identifying what the correct white balance should have been, but in the fact that the in-camera white-balance procedure is followed by several camera-specific nonlinear color manipulations that make it challenging to correct the image's colors in post-processing. This paper introduces the first method to explicitly address this problem. Our method is enabled by a dataset of over 65,000 pairs of incorrectly white-balanced images and their corresponding correctly white-balanced images. Using this dataset, we introduce a k-nearest neighbor strategy that is able to compute a nonlinear color mapping function to correct the image's colors. We show our method is highly effective and generalizes well to camera models not in the training set. Mahmoud Afifi, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 3 |
| 2019 | MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User InputabstractExisting deep learning-based interactive image segmentation approaches typically assume the target-of-interest is always a single object and fail to account for the potential diversity in user expectations, thus requiring excessive user input when it comes to segmenting an object part or a group of objects instead. Motivated by the observation that the object part, full object, and a collection of objects essentially differ in size, we propose a new concept called scale-diversity, which characterizes the spectrum of segmentations w.r.t. different scales. To address this, we present MultiSeg, a scale-diverse interactive image segmentation network that incorporates a set of two-dimensional scale priors into the model to generate a set of scale-varying proposals that conform to the user input. We explicitly encourage segmentation diversity during training by synthesizing diverse training samples for a given image. As a result, our method allows the user to quickly locate the closest segmentation target for further refinement if necessary. Despite its simplicity, experimental results demonstrate that our proposed model is capable of quickly producing diverse yet plausible segmentation outputs, reducing the user interaction required, especially in cases where many types of segmentations (object parts or groups) are expected. Jun Hao Liew, Scott Cohen, Brian L. Price, Long Mai, Sim Heng Ong, Jiashi Feng |
ICCV | 2 |
| 2019 | Unconstrained Foreground Object SearchabstractMany people search for foreground objects to use when editing images. While existing methods can retrieve candidates to aid in this, they are constrained to returning objects that belong to a pre-specified semantic class. We instead propose a novel problem of unconstrained foreground object (UFO) search and introduce a solution that supports efficient search by encoding the background image in the same latent space as the candidate foreground objects. A key contribution of our work is a cost-free, scalable approach for creating a large-scale training dataset with a variety of foreground objects of differing semantic categories per image location. Quantitative and human-perception experiments with two diverse datasets demonstrate the advantage of our UFO search solution over related baselines. Brian L. Price, Scott Cohen, Danna Gurari |
ICCV | 3 |
| 2019 | Deep Visual Template-Free Form ParsingabstractThe following topics are dealt with: learning (artificial intelligence); document image processing; feature extraction; text analysis; convolutional neural nets; image segmentation; handwritten character recognition; image classification; text detection; optical character recognition. Brian L. Davis, Bryan S. Morse, Scott Cohen, Brian L. Price, Chris Tensmeyer |
ICDAR | 3 |
| 2019 | Deep Splitting and Merging for Table Structure DecompositionabstractGiven the large variety and complexity of tables, table structure extraction is a challenging task in automated document analysis systems. We present a pair of novel deep learning models (Split and Merge models) that given an input image, 1) predicts the basic table grid pattern and 2) predicts which grid elements should be merged to recover cells that span multiple rows or columns. We propose projection pooling as a novel component of the Split model and grid pooling as a novel part of the Merge model. While most Fully Convolutional Networks rely on local evidence, these unique pooling regions allow our models to take advantage of the global table structure. We achieve state-of-the-art performance on the public ICDAR 2013 Table Competition dataset of PDF documents. On a much larger private dataset which we used to train the models, we significantly outperform both a state-ofthe-art deep model and a major commercial software system. Chris Tensmeyer, Vlad I. Morariu, Brian L. Price, Scott Cohen, Tony R. Martinez |
ICDAR | 4 |
| 2019 | Multi-label Connectionist Temporal ClassificationabstractThe Connectionist Temporal Classification (CTC) loss function [1] enables end-to-end training of a neural network for sequence-to-sequence tasks without the need for prior alignments between the input and output. CTC is traditionally used for training sequential, single-label problems; each element in the sequence has only one class. In this work, we show that CTC is not suitable for multi-label tasks and we present a novel Multi-label Connectionist Temporal Classification (MCTC) loss function for multi-label, sequence-to-sequence classification. Multi-label classes can represent meaningful attributes of a single element; for example, in Optical Music Recognition (OMR), a music note can have separate duration and pitch attributes. Our approach achieves state-of-the-art results on Joint Handwritten Text Recognition and Name Entity Recognition, Asian Character Recognition, and OMR. Curtis Wigington, Brian L. Price, Scott Cohen |
ICDAR | 3 |
| 2019 | Guided Image Inpainting: Replacing an Image Region by Pulling Content From Another ImageabstractDeep generative models have shown success in automatically synthesizing missing image regions using surrounding context. However, users cannot directly decide what content to synthesize with such approaches.We propose an end-to-end network for image inpainting that uses a different image to guide the synthesis of new content to fill the hole. A key challenge addressed by our approach is synthesizing new content in regions where the guidance image and the context of the original image are inconsistent. We conduct four studies that demonstrate our method yields more realistic image inpainting results over seven baselines. Brian L. Price, Scott Cohen, Danna Gurari |
WACV | 3 |
| 2018 | Progressive Attention Networks for Visual Attribute Prediction
Hongsuck Seo, Zhe Lin 0001, Scott Cohen, Xiaohui Shen, Bohyung Han |
BMVC | 3 |
| 2018 | DVQA: Understanding Data Visualizations via Question AnsweringabstractBar charts are an effective way to convey numeric information, but today's algorithms cannot parse them. Existing methods fail when faced with even minor variations in appearance. Here, we present DVQA, a dataset that tests many aspects of bar chart understanding in a question answering framework. Unlike visual question answering (VQA), DVQA requires processing words and answers that are unique to a particular bar chart. State-of-the-art VQA algorithms perform poorly on DVQA, and we propose two strong baselines that perform considerably better. Our work will enable algorithms to automatically extract numeric and semantic information from vast quantities of bar charts found in scientific publications, Internet articles, business reports, and many other areas. Kushal Kafle, Brian L. Price, Scott Cohen, Christopher Kanan |
CVPR | 3 |
| 2018 | Discriminability Objective for Training Descriptive CaptionsabstractOne property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning. Ruotian Luo, Brian L. Price, Scott Cohen, Gregory Shakhnarovich |
CVPR | 3 |
| 2018 | Interactive Boundary Prediction for Object Selection
Hoang Le, Long Mai, Brian L. Price, Scott Cohen, Hailin Jin, Feng Liu 0015 |
ECCV (14) | 4 |
| 2018 | Concept Mask: Large-Scale Segmentation from Semantic Concepts
Yufei Wang 0001, Zhe Lin 0001, Xiaohui Shen, Jianming Zhang 0001, Scott Cohen |
ECCV (12) | 5 |
| 2018 | Start, Follow, Read: End-to-End Full-Page Handwriting Recognition
Curtis Wigington, Chris Tensmeyer, Brian L. Davis, Bill Barrett, Brian L. Price, Scott Cohen |
ECCV (6) | 6 |
| 2018 | YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
Ning Xu 0007, Yuchen Fan 0001, Jianchao Yang, Dingcheng Yue, Brian L. Price, Scott Cohen, Thomas S. Huang |
ECCV (5) | 8 |
| 2017 | Sherlock: Scalable Fact Learning in ImagesabstractWe study scalable and uniform understanding of facts in images. Existing visual recognition systems are typically modeled differently for each fact type such as objects, actions, and interactions. We propose a setting where all these facts can be modeled simultaneously with a capacity to understand an unbounded number of facts in a structured way. The training data comes as structured facts in images, including (1) objects (e.g., ), (2) attributes (e.g., ), (3) actions (e.g., ), and (4) interactions (e.g., ). Each fact has a semantic language view (e.g., < boy, playing>) and a visual view (an image with this fact). We show that learning visual facts in a structured way enables not only a uniform but also generalizable visual understanding. We propose and investigate recent and strong approaches from the multiview learning literature and also introduce two learning representation models as potential baselines. We applied the investigated methods on several datasets that we augmented with structured facts and a large scale dataset of more than 202,000 facts and 814,000 images. Our experiments show the advantage of relating facts by the structure by the proposed models compared to the designed baselines on bidirectional fact retrieval. Scott Cohen, Walter Chang, Brian L. Price, Ahmed M. Elgammal |
AAAI | 2 |
| 2017 | Deep GrabCut for Object Selection
Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang |
BMVC | 3 |
| 2017 | Forecasting Human Dynamics from Static ImagesabstractThis paper presents the first study on forecasting human dynamics from static images. The problem is to input a single RGB image and generate a sequence of upcoming human body poses in 3D. To address the problem, we propose the 3D Pose Forecasting Network (3D-PFNet). Our 3D-PFNet integrates recent advances on single-image human pose estimation and sequence prediction, and converts the 2D predictions into 3D space. We train our 3D-PFNet using a three-step training strategy to leverage a diverse source of training data, including image and video based human pose datasets and 3D motion capture (MoCap) data. We demonstrate competitive performance of our 3D-PFNet on 2D pose forecasting and 3D structure recovery through quantitative and qualitative results. Yu-Wei Chao, Jimei Yang, Brian L. Price, Scott Cohen, Jia Deng 0001 |
CVPR | 4 |
| 2017 | Depth from Defocus in the WildabstractWe consider the problem of two-frame depth from defocus in conditions unsuitable for existing methods yet typical of everyday photography: a non-stationary scene, a handheld cellphone camera, a small aperture, and sparse scene texture. The key idea of our approach is to combine local estimation of depth and flow in very small patches with a global analysis of image content-3D surfaces, deformations, figure-ground relations, textures. To enable local estimation we (1) derive novel defocus-equalization filters that induce brightness constancy across frames and (2) impose a tight upper bound on defocus blur-just three pixels in radius-by appropriately refocusing the camera for the second input frame. For global analysis we use a novel splinebased scene representation that can propagate depth and flow across large irregularly-shaped regions. Our experiments show that this combination preserves sharp boundaries and yields good depth and flow maps in the face of significant noise, non-rigidity, and data sparsity. Huixuan Tang, Scott Cohen, Brian L. Price, Stephen Schiller, Kiriakos N. Kutulakos |
CVPR | 2 |
| 2017 | Skeleton Key: Image Captioning by Skeleton-Attribute DecompositionabstractRecently, there has been a lot of interest in automatically generating descriptions for an image. Most existing language-model based approaches for this task learn to generate an image description word by word in its original word order. However, for humans, it is more natural to locate the objects and their relationships first, and then elaborate on each object, describing notable attributes. We present a coarse-to-fine method that decomposes the original image description into a skeleton sentence and its attributes, and generates the skeleton sentence and attribute phrases separately. By this decomposition, our method can generate more accurate and novel descriptions than the previous state-of-the-art. Experimental results on the MS-COCO and a larger scale Stock3M datasets show that our algorithm yields consistent improvements across different evaluation metrics, especially on the SPICE metric, which has much higher correlation with human ratings than the conventional metrics. Furthermore, our algorithm can generate descriptions with varied length, benefiting from the separate control of the skeleton and attributes. This enables image description generation that better accommodates user preferences. Yufei Wang 0001, Zhe Lin 0001, Xiaohui Shen, Scott Cohen, Garrison W. Cottrell |
CVPR | 4 |
| 2017 | Deep Image MattingabstractImage matting is a fundamental computer vision problem and has many applications. Previous algorithms have poor performance when an image has similar foreground and background colors or complicated textures. The main reasons are prior methods 1) only use low-level features and 2) lack high-level context. In this paper, we propose a novel deep learning based algorithm that can tackle both these problems. Our deep model has two parts. The first part is a deep convolutional encoder-decoder network that takes an image and the corresponding trimap as inputs and predict the alpha matte of the image. The second part is a small convolutional network that refines the alpha matte predictions of the first network to have more accurate alpha values and sharper edges. In addition, we also create a large-scale image matting dataset including 49300 training images and 1000 testing images. We evaluate our algorithm on the image matting benchmark, our testing set, and a wide variety of real images. Experimental results clearly demonstrate the superiority of our algorithm over previous methods. Ning Xu 0007, Brian L. Price, Scott Cohen, Thomas S. Huang |
CVPR | 3 |
| 2017 | Relationship Proposal NetworksabstractImage scene understanding requires learning the relationships between objects in the scene. A scene with many objects may have only a few individual interacting objects (e.g., in a party image with many people, only a handful of people might be speaking with each other). To detect all relationships, it would be inefficient to first detect all individual objects and then classify all pairs, not only is the number of all pairs quadratic, but classification requires limited object categories, which is not scalable for real-world images. In this paper we address these challenges by using pairs of related regions in images to train a relationship proposer that at test time produces a manageable number of related regions. We name our model the Relationship Proposal Network (Rel-PN). Like object proposals, our Rel-PN is class-agnostic and thus scalable to an open vocabulary of objects. We demonstrate the ability of our Rel-PN to localize relationships with only a few thousand proposals. We demonstrate its performance on the Visual Genome dataset and compare to other baselines that we designed. We also conduct experiments on a smaller subset of 5,000 images with over 37,000 related regions and show promising results. Scott Cohen, Walter Chang, Ahmed M. Elgammal |
CVPR | 3 |
| 2017 | Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table DetectionabstractPage segmentation and table detection play an important role in understanding the structure of documents. We present a page segmentation algorithm that incorporates state-of-the-art deep learning methods for segmenting three types of document elements: text blocks, tables, and figures. We propose a multi-scale, multi-task fully convolutional neural network (FCN) for the tasks of semantic page segmentation and element contour detection. The semantic segmentation network accurately predicts the probability at each pixel of the three element classes. The contour detection network accurately predicts instance level "edges" around each element occurrence. We propose a conditional random field (CRF) that uses features output from the semantic segmentation and contour networks to improve upon the semantic segmentation network output. Given the semantic segmentation output, we also extract individual table instances from the page using some heuristic rules and a verification network to remove false positives. We show that although we only consider a page image as input, we produce comparable results with other methods that relies on PDF file information and heuristics and hand crafted features tailored to specific types of documents. Our approach learns the representative features for page segmentation from real and synthetic training data. %, and produces good results on real documents. The learning-based property makes it a more general method than existing methods in terms of document types and element appearances. For example, our method reliably detects sparsely lined tables which are hard for rule-based or heuristic methods. Dafang He, Scott Cohen, Brian L. Price, Daniel Kifer, C. Lee Giles |
ICDAR | 2 |
| 2017 | Data Augmentation for Recognition of Handwritten Words and Lines Using a CNN-LSTM NetworkabstractWe introduce two data augmentation and normalization techniques, which, used with a CNN-LSTM, significantly reduce Word Error Rate (WER) and Character Error Rate (CER) beyond best-reported results on handwriting recognition tasks. (1) We apply a novel profile normalization technique to both word and line images. (2) We augment existing text images using random perturbations on a regular grid. We apply our normalization and augmentation to both training and test images. Our approach achieves low WER and CER over hundreds of authors, multiple languages and a variety of collections written centuries apart. Image augmentation in this manner achieves state-of-the-art recognition accuracy on several popular handwritten word benchmarks. Curtis Wigington, Seth Stewart, Brian L. Davis, Bill Barrett, Brian L. Price, Scott Cohen |
ICDAR | 6 |
| 2017 | Group-Theme Recoloring for Multi-Image Color ConsistencyabstractAbstract Modifying the colors of an image is a fundamental editing task with a wide range of methods available. Manipulating multiple images to share similar colors is more challenging, with limited tools available. Methods such as color transfer are effective in making an image share similar colors with a target image; however, color transfer is not suitable for modifying multiple images. Approaches for color consistency for photo collections give good results when the photo collection contains similar scene content, but are not applicable for general input images. To address these gaps, we propose an application framework for achieving color consistency for multi‐image input. Our framework derives a group color theme from the input images′ individual color palettes and uses this group color theme to recolor the image collection. This group‐theme recoloring provides an effective way to ensure color consistency among multiple images and naturally lends itself to the inclusion of an additional external color theme. We detail our group‐theme recoloring approach and demonstrate its effectiveness on a number of examples. Nguyen Ho Man Rang, Brian L. Price, Scott Cohen, Michael S. Brown |
Comput. Graph. Forum | 3 |
| 2016 | Two Illuminant Estimation and User Correction PreferenceabstractThis paper examines the problem of white-balance correction when a scene contains two illuminations. This is a two step process: 1) estimate the two illuminants, and 2) correct the image. Existing methods attempt to estimate a spatially varying illumination map, however, results are error prone and the resulting illumination maps are too lowresolution to be used for proper spatially varying whitebalance correction. In addition, the spatially varying nature of these methods make them computationally intensive. We show that this problem can be effectively addressed by not attempting to obtain a spatially varying illumination map, but instead by performing illumination estimation on large sub-regions of the image. Our approach is able to detect when distinct illuminations are present in the image and accurately measure these illuminants. Since our proposed strategy is not suitable for spatially varying image correction, a user study is performed to see if there is a preference for how the image should be corrected when two illuminants are present, but only a global correction can be applied. The user study shows that when the illuminations are distinct, there is a preference for the outdoor illumination to be corrected resulting in warmer final result. We use these collective findings to demonstrate an effective two illuminant estimation scheme that produces corrected images that users prefer. Dongliang Cheng, Abdelrahman Kamel, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 4 |
| 2016 | Interactive Segmentation on RGBD Images via Cue SelectionabstractInteractive image segmentation is an important problem in computer vision with many applications including image editing, object recognition and image retrieval. Most existing interactive segmentation methods only operate on color images. Until recently, very few works have been proposed to leverage depth information from low-cost sensors to improve interactive segmentation. While these methods achieve better results than color-based methods, they are still limited in either using depth as an additional color channel or simply combining depth with color in a linear way. We propose a novel interactive segmentation algorithm which can incorporate multiple feature cues like color, depth, and normals in an unified graph cut framework to leverage these cues more effectively. A key contribution of our method is that it automatically selects a single cue to be used at each pixel, based on the intuition that only one cue is necessary to determine the segmentation label locally. This is achieved by optimizing over both segmentation labels and cue labels, using terms designed to decide where both the segmentation and label cues should change. Our algorithm thus produces not only the segmentation mask but also a cue label map that indicates where each cue contributes to the final result. Extensive experiments on five large scale RGBD datasets show that our proposed algorithm performs significantly better than both other color-based and RGBD based algorithms in reducing the amount of user inputs as well as increasing segmentation accuracy. Brian L. Price, Scott Cohen, Shih-Fu Chang |
CVPR | 3 |
| 2016 | Deep Interactive Object SelectionabstractInteractive object selection is a very important research problem and has many applications. Previous algorithms require substantial user interactions to estimate the foreground and background distributions. In this paper, we present a novel deep-learning-based algorithm which has much better understanding of objectness and can reduce user interactions to just a few clicks. Our algorithm transforms user-provided positive and negative clicks into two Euclidean distance maps which are then concatenated with the RGB channels of images to compose (image, user interactions) pairs. We generate many of such pairs by combining several random sampling strategies to model users' click patterns and use them to finetune deep Fully Convolutional Networks (FCNs). Finally the output probability maps of our FCN-8s model is integrated with graph cut optimization to refine the boundary segments. Our model is trained on the PASCAL segmentation dataset and evaluated on other datasets with different object classes. Experimental results on both seen and unseen objects demonstrate that our algorithm has a good generalization ability and is superior to all existing interactive object selection approaches. Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang |
CVPR | 3 |
| 2016 | Object Contour Detection with a Fully Convolutional Encoder-Decoder NetworkabstractWe develop a deep learning algorithm for contour detection with a fully convolutional encoder-decoder network. Different from previous low-level edge detection, our algorithm focuses on detecting higher-level object contours. Our network is trained end-to-end on PASCAL VOC with refined ground truth from inaccurate polygon annotations, yielding much higher precision in object contour detection than previous methods. We find that the learned model generalizes well to unseen object classes from the same supercategories on MS COCO and can match state-of-the-art edge detection on BSDS500 with fine-tuning. By combining with the multiscale combinatorial grouping algorithm, our method can generate high-quality segmented object proposals, which significantly advance the state-of-the-art on PASCAL VOC (improving average recall from 0.62 to 0.67) with a relatively small amount of candidates (~1660 per image). Jimei Yang, Brian L. Price, Scott Cohen, Honglak Lee, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2016 | SURGE: Surface Regularized Geometry Estimation from a Single ImageabstractThis paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all while maintaining fine detail within objects. Our approach comprises two components: (i) a fourstream convolutional neural network (CNN) where depths, surface normals, and likelihoods of planar region and planar boundary are predicted at each pixel, followed by (ii) a dense conditional random field (DCRF) that integrates the four predictions such that the normals and depths are compatible with each other and regularized by the planar region and planar boundary information. The DCRF is formulated such that gradients can be passed to the surface normal and depth CNNs via backpropagation. In addition, we propose new planar wise metrics to evaluate geometry consistency within planar surfaces, which are more tightly related to dependent 3D editing applications. We show that our regularization yields a 30% relative improvement in planar consistency on the NYU v2 dataset. Peng Wang 0001, Xiaohui Shen, Bryan C. Russell, Scott Cohen, Brian L. Price, Alan L. Yuille |
NIPS | 4 |
| 2015 | Effective learning-based illuminant estimation using simple featuresabstractIllumination estimation is the process of determining the chromaticity of the illumination in an imaged scene in order to remove undesirable color casts through white-balancing. While computational color constancy is a well-studied topic in computer vision, it remains challenging due to the ill-posed nature of the problem. One class of techniques relies on low-level statistical information in the image color distribution and works under various assumptions (e.g. Grey-World, White-Patch, etc). These methods have an advantage that they are simple and fast, but often do not perform well. More recent state-of-the-art methods employ learning-based techniques that produce better results, but often rely on complex features and have long evaluation and training times. In this paper, we present a learning-based method based on four simple color features and show how to use this with an ensemble of regression trees to estimate the illumination. We demonstrate that our approach is not only faster than existing learning-based methods in terms of both evaluation and training time, but also gives the best results reported to date on modern color constancy data sets. Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown |
CVPR | 3 |
| 2015 | Towards unified depth and semantic prediction from a single imageabstractDepth estimation and semantic segmentation are two fundamental problems in image understanding. While the two tasks are strongly correlated and mutually beneficial, they are usually solved separately or sequentially. Motivated by the complementary properties of the two tasks, we propose a unified framework for joint depth and semantic prediction. Given an image, we first use a trained Convolutional Neural Network (CNN) to jointly predict a global layout composed of pixel-wise depth values and semantic labels. By allowing for interactions between the depth and semantic information, the joint network provides more accurate depth prediction than a state-of-the-art CNN trained solely for depth prediction [6]. To further obtain fine-level details, the image is decomposed into local segments for region-level depth and semantic prediction under the guidance of global layout. Utilizing the pixel-wise global prediction and region-wise local prediction, we formulate the inference problem in a two-layer Hierarchical Conditional Random Field (HCRF) to produce the final depth and semantic map. As demonstrated in the experiments, our approach effectively leverages the advantages of both tasks and provides the state-of-the-art results. Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille |
CVPR | 4 |
| 2015 | PatchCut: Data-driven object segmentation via local shape transferabstractObject segmentation is highly desirable for image understanding and editing. Current interactive tools require a great deal of user effort while automatic methods are usually limited to images of special object categories or with high color contrast. In this paper, we propose a data-driven algorithm that uses examples to break through these limits. As similar objects tend to share similar local shapes, we match query image patches with example images in multiscale to enable local shape transfer. The transferred local shape masks constitute a patch-level segmentation solution space and we thus develop a novel cascade algorithm, PatchCut, for coarse-to-fine object segmentation. In each stage of the cascade, local shape mask candidates are selected to refine the estimated segmentation of the previous stage iteratively with color models. Experimental results on various datasets (Weizmann Horse, Fashionista, Object Discovery and PASCAL) demonstrate the effectiveness and robustness of our algorithm. Jimei Yang, Brian L. Price, Scott Cohen, Zhe Lin 0001, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2015 | Beyond White: Ground Truth Colors for Color Constancy CorrectionabstractA limitation in color constancy research is the inability to establish ground truth colors for evaluating corrected images. Many existing datasets contain images of scenes with a color chart included, however, only the chart's neutral colors (grayscale patches) are used to provide the ground truth for illumination estimation and correction. This is because the corrected neutral colors are known to lie along the achromatic line in the camera's color space (i.e. R=G=B), the correct RGB values of the other color patches are not known. As a result, most methods estimate a 3*3 diagonal matrix that ensures only the neutral colors are correct. In this paper, we describe how to overcome this limitation. Specifically, we show that under certain illuminations, a diagonal 3*3 matrix is capable of correcting not only neutral colors, but all the colors in a scene. This finding allows us to find the ground truth RGB values for the color chart in the camera's color space. We show how to use this information to correct all the images in existing datasets to have correct colors. Working from these new color corrected datasets, we describe how to modify existing color constancy algorithms to perform better image correction. Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown |
ICCV | 3 |
| 2015 | Joint Object and Part Segmentation Using Deep Learned PotentialsabstractSegmenting semantic objects from images and parsing them into their respective semantic parts are fundamental steps towards detailed object understanding in computer vision. In this paper, we propose a joint solution that tackles semantic object and part segmentation simultaneously, in which higher object-level context is provided to guide part segmentation, and more detailed part-level localization is utilized to refine object segmentation. Specifically, we first introduce the concept of semantic compositional parts (SCP) in which similar semantic parts are grouped and shared among different objects. A two-stream fully convolutional network (FCN) is then trained to provide the SCP and object potentials at each pixel. At the same time, a compact set of segments can also be obtained from the SCP predictions of the network. Given the potentials and the generated segments, in order to explore long-range context, we finally construct an efficient fully connected conditional random field (FCRF) to jointly predict the final object and part labels. Extensive evaluation on three different datasets shows that our approach can mutually enhance the performance of object and part segmentation, and outperforms the current state-of-the-art on both tasks. Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille |
ICCV | 4 |
| 2014 | Semantic Object SelectionabstractInteractive object segmentation has great practical importance in computer vision. Many interactive methods have been proposed utilizing user input in the form of mouse clicks and mouse strokes, and often requiring a lot of user intervention. In this paper, we present a system with a far simpler input method: the user needs only give the name of the desired object. With the tag provided by the user we do a text query of an image database to gather exemplars of the object. Using object proposals and borrowing ideas from image retrieval and object detection, the object is localized in the target image. An appearance model generated from the exemplars and the location prior are used in an energy minimization framework to select the object. Our method outperforms the state-of-the-art on existing datasets and on a more challenging dataset we collected. Ejaz Ahmed 0002, Scott Cohen, Brian L. Price |
CVPR | 2 |
| 2014 | Context Driven Scene Parsing with Attention to Rare ClassesabstractThis paper presents a scalable scene parsing algorithm based on image retrieval and superpixel matching. We focus on rare object classes, which play an important role in achieving richer semantic understanding of visual scenes, compared to common background classes. Towards this end, we make two novel contributions: rare class expansion and semantic context description. First, considering the long-tailed nature of the label distribution, we expand the retrieval set by rare class exemplars and thus achieve more balanced superpixel classification results. Second, we incorporate both global and local semantic context information through a feedback based mechanism to refine image retrieval and superpixel matching. Results on the SIFTflow and LMSun datasets show the superior performance of our algorithm, especially on the rare classes, without sacrificing overall labeling accuracy. Jimei Yang, Brian L. Price, Scott Cohen, Ming-Hsuan Yang 0001 |
CVPR | 3 |
| 2014 | Depth-based patch scaling for content-aware stereo image completionabstractA number of recent algorithms have been proposed for working with stereo image pairs in ways that are already familiar to users of single-image editing tools. In particular, Morse, et al. (2012) have proposed a method for performing image completion in stereo images so as to maintain stereoscopic consistency. Like prior work in stereo completion, this method drew source texture only from regions at the same depth as the target region, which while helping the result can sometimes overly limit the pool of suitable source textures. Other methods such as the Generalized PatchMatch approach of Barnes, et al. (2010) have used scaled (and otherwise transformed) source texture to improve the quality of the completed target region, but these methods rely on randomly sampling the scale (or transformation) space without knowledge of scene geometry. This paper extends stereo image completion to include source textures scaled according to the relative differences in depth between image regions. Limited random sampling is used to make the method robust to minor errors in the stereo disparities and to provide for non-uniform aspect ratios, but with far fewer random samples than prior unrestrained sampling of scale. A preference for unscaled or downsampled source textures rather than upsampled ones is incorporated into the objective function and avoids an inherent matching bias towards low-frequency regions. Results demonstrate that using scene geometry to drive scale selection results in improved image completion compared to either single-image completion or prior methods for stereo completion. Joel Howard, Bryan S. Morse, Scott Cohen, Brian L. Price |
WACV | 3 |
| 2014 | Temporally coherent and spatially accurate video mattingabstractAbstract Image and video matting are still challenging problems in areas with low foreground‐background contrast. Video matting also has the challenge of ensuring temporally coherent mattes because the human visual system is highly sensitive to temporal jitter and flickering. On the other hand, video provides the opportunity to use information from other frames to improve the matte accuracy on a given frame. In this paper, we present a new video matting approach that improves the temporal coherence while maintaining high spatial accuracy in the computed mattes. We build sample sets of temporal and local samples that cover all the color distributions of the object and background over all previous frames. This helps guarantee spatial accuracy and temporal coherence by ensuring that proper samples are found even when distantly located in space or time. An explicit energy term encourages temporal consistency in the mattes derived from the selected samples. In addition, we use localized texture features to improve spatial accuracy in low contrast regions where color distributions overlap. The proposed method results in better spatial accuracy and temporal coherence than existing video matting methods. E. Shahrian, Brian L. Price, Scott Cohen, D. Rajan |
Comput. Graph. Forum | 3 |
| 2013 | Stereo+Kinect for High Resolution Stereo CorrespondencesabstractIn this work, we combine the complementary depth sensors Kinect and stereo image matching to obtain high quality correspondences. Our goal is to obtain a dense disparity map at the spatial and depth resolution of the stereo cameras (4-12 MP). We propose a global optimization scheme, where both the data and smoothness costs are derived using sensor confidences and low resolution geometry from Kinect. A spatially varying search range is used to limit the number of potential disparities at each pixel. The smoothness prior is Based on available low resolution depth from Kinect rather than image gradients, thus performing better in both textured areas with smooth depth and texture-less areas with depth gradient. We also propose a spatially varying smoothness weight to better handle occlusion areas, and the relative contribution of the two energy terms. We demonstrate how the two sensors can be effectively fused to obtain correct scene depth in ambiguous areas, as well as fine structural details in textured areas. Gowri Somanath, Scott Cohen, Brian L. Price, Chandra Kambhamettu |
3DV | 2 |
| 2013 | High-Quality Stereo Video Matching via User Interaction and Space-Time PropagationabstractEven current state-of-the-art automatic stereo matching methods often struggle on natural images and videos, in great part due to fundamental matching ambiguities in low texture regions and a lack of higher level object knowledge. Stereo image matching can benefit greatly from user input to guide the matching process and help disambiguate matches. Applying interactive correction tools from scratch on each frame of a video would not only be throwing away valuable information provided by the user on other frames, but would also likely be too time consuming to be practical for video even if excellent disparity results could be obtained within a few minutes on each frame. In this work, we propose a stereo video matching system that allows user interaction to obtain high quality, dense disparity maps on key frames and then intelligently propagates the user input and key frame disparities to automatically produce high quality disparity maps on intermediate frames. The disparity maps on key frames are obtained using several novel, easy-to-use, and effective interactive tools. Our novel propagation algorithm estimates 3D transformations that map user corrected areas in key frames to intermediate frames. Experiments demonstrate the effectiveness and efficiency of our hybrid interactive/automatic approach. Brian L. Price, Scott Cohen, Ruigang Yang |
3DV | 3 |
| 2013 | Large Displacement Optical Flow from Nearest Neighbor FieldsabstractWe present an optical flow algorithm for large displacement motions. Most existing optical flow methods use the standard coarse-to-fine framework to deal with large displacement motions which has intrinsic limitations. Instead, we formulate the motion estimation problem as a motion segmentation problem. We use approximate nearest neighbor fields to compute an initial motion field and use a robust algorithm to compute a set of similarity transformations as the motion candidates for segmentation. To account for deviations from similarity transformations, we add local deformations in the segmentation process. We also observe that small objects can be better recovered using translations as the motion candidates. We fuse the motion results obtained under similarity transformations and under translations together before a final refinement. Experimental validation shows that our method can successfully handle large displacement motions. Although we particularly focus on large displacement motions in this work, we make no sacrifice in terms of overall performance. In particular, our method ranks at the top of the Middlebury benchmark. Zhuoyuan Chen, Hailin Jin, Zhe Lin 0001, Scott Cohen, Ying Wu 0001 |
CVPR | 4 |
| 2013 | Improving Image Matting Using Comprehensive Sampling SetsabstractIn this paper, we present a new image matting algorithm that achieves state-of-the-art performance on a benchmark dataset of images. This is achieved by solving two major problems encountered by current sampling based algorithms. The first is that the range in which the foreground and background are sampled is often limited to such an extent that the true foreground and background colors are not present. Here, we describe a method by which a more comprehensive and representative set of samples is collected so as not to miss out on the true samples. This is accomplished by expanding the sampling range for pixels farther from the foreground or background boundary and ensuring that samples from each color distribution are included. The second problem is the overlap in color distributions of foreground and background regions. This causes sampling based methods to fail to pick the correct samples for foreground and background. Our design of an objective function forces those foreground and background samples to be picked that are generated from well-separated distributions. Comparison on the dataset at and evaluation by www.alphamatting.com shows that the proposed method ranks first in terms of error measures used in the website. Ehsan Shahrian, Deepu Rajan, Brian L. Price, Scott Cohen |
CVPR | 4 |
| 2013 | Fast Image Super-Resolution Based on In-Place Example RegressionabstractWe propose a fast regression model for practical single image super-resolution based on in-place examples, by leveraging two fundamental super-resolution approaches- learning from an external database and learning from self-examples. Our in-place self-similarity refines the recently proposed local self-similarity by proving that a patch in the upper scale image have good matches around its origin location in the lower scale image. Based on the in-place examples, a first-order approximation of the nonlinear mapping function from low-to high-resolution image patches is learned. Extensive experiments on benchmark and real-world images demonstrate that our algorithm can produce natural-looking results with sharp edges and preserved fine details, while the current state-of-the-art algorithms are prone to visual artifacts. Furthermore, our model can easily extend to deal with noise by combining the regression results on multiple in-place examples for robust estimation. The algorithm runs fast and is particularly useful for practical applications, where the input images typically contain diverse textures and they are potentially contaminated by noise or compression artifacts. Jianchao Yang, Zhe Lin 0001, Scott Cohen |
CVPR | 3 |
| 2013 | Estimating Spatially Varying Defocus Blur From A Single ImageabstractEstimating the amount of blur in a given image is important for computer vision applications. More specifically, the spatially varying defocus point-spread-functions (PSFs) over an image reveal geometric information of the scene, and their estimate can also be used to recover an all-in-focus image. A PSF for a defocus blur can be specified by a single parameter indicating its scale. Most existing algorithms can only select an optimal blur from a finite set of candidate PSFs for each pixel. Some of those methods require a coded aperture filter inserted in the camera. In this paper, we present an algorithm estimating a defocus scale map from a single image, which is applicable to conventional cameras. This method is capable of measuring the probability of local defocus scale in the continuous domain. It also takes smoothness and color edge information into consideration to generate a coherent blur map indicating the amount of blur at each pixel. Simulated and real data experiments illustrate excellent performance and its successful applications in foreground/background segmentation. Scott Cohen, Stephen Schiller, Peyman Milanfar |
IEEE Trans. Image Process. | 2 |
| 2012 | Video upscaling via spatio-temporal self-similarity
Alper Ayvaci, Hailin Jin, Zhe Lin 0001, Scott Cohen, Stefano Soatto |
ICPR | 4 |
| 2012 | Coupled Dictionary Training for Image Super-ResolutionabstractIn this paper, we propose a novel coupled dictionary training method for single image super-resolution based on patchwise sparse recovery, where the learned couple dictionaries relate the low- and high-resolution image patch spaces via sparse representation. The learning process enforces that the sparse representation of a low-resolution image patch in terms of the low-resolution dictionary can well reconstruct its underlying high-resolution image patch with the dictionary in the highresolution image patch space. We model the learning problem as a bilevel optimization problem, where the optimization includes an 1-norm minimization problem in its constraints. Implicit differentiation is employed to calculate the desired gradient for stochastic gradient descent. We demonstrate that our coupled dictionary learning method can outperform the existing joint dictionary training method both quantitatively and qualitatively. Furthermore, for real applications, we speed up the algorithm approximately 10 times by learning a neural network model for fast sparse inference and selectively processing only those visually salient regions. Extensive experimental comparisons with stateof- the-art super-resolution algorithms validate the effectiveness of our proposed approach. Jianchao Yang, Zhe Lin 0001, Scott Cohen, Thomas S. Huang |
IEEE Trans. Image Process. | 4 |
| 2011 | StereoCut: Consistent interactive object selection in stereo image pairsabstractMethods of interacting with stereo image pairs are important for handling the increasing amount of stereoscopic 3D data now being produced. In this paper, we introduce a framework for interactively selecting objects in two stereo images simultaneously using graph cut. A key contribution of our method is the use of stereo correspondence probability distributions to govern the strength of the connection between the two images. This allows information from arbitrary stereo matching algorithms to be utilized by our method. We show how to enforce consistency in these distributions to improve the results. For comparisons, we introduce a new dataset of stereo images and ground truth selections. We evaluate different correspondence distributions and show that our method is effective in selecting objects from stereo pairs. Brian L. Price, Scott Cohen |
ICCV | 2 |
| 2010 | Simultaneous foreground, background, and alpha estimation for image mattingabstractImage matting is the process of extracting a soft segmentation of an object in an image as defined by the matting equation. Most current techniques focus largely on computing the alpha values of unknown pixels and treat computation of the foreground and background colors as an afterthought, if at all. However, for many applications, such as compositing an object into a new scene or deleting an object from the scene, the foreground and background colors are vital for an acceptable answer. We propose a method of solving for the foreground, background, and alpha of an unknown region in an image simultaneously. This allows for novel constraints to be placed directly on the foreground and background as well as on alpha. We show through both visual results and quantitative measurements on standard datasets that this approach produces more accurate foreground and background values at each pixel while maintaining competitive results on the alpha matte. Brian L. Price, Bryan S. Morse, Scott Cohen |
CVPR | 3 |
| 2010 | Geodesic graph cut for interactive image segmentationabstractInteractive segmentation is useful for selecting objects of interest in images and continues to be a topic of much study. Methods that grow regions from foreground/background seeds, such as the recent geodesic segmentation approach, avoid the boundary-length bias of graph-cut methods but have their own bias towards minimizing paths to the seeds, resulting in increased sensitivity to seed placement. The lack of edge modeling in geodesic or similar approaches limits their ability to precisely localize object boundaries, something at which graph-cut methods generally excel. This paper presents a method for combining geodesic-distance information with edge information in a graphcut optimization framework, leveraging the complementary strengths of each. Rather than a fixed combination we use the distinctiveness of the foreground/background color models to predict the effectiveness of the geodesic distance term and adjust the weighting accordingly. We also introduce a spatially varying weighting that decreases the potential for shortcutting in object interiors while transferring greater control to the edge term for better localization near object boundaries. Results show our method is less prone to shortcutting than typical graph cut methods while being less sensitive to seed placement and better at edge localization than geodesic methods. This leads to increased segmentation accuracy and reduced effort on the part of the user. Brian L. Price, Bryan S. Morse, Scott Cohen |
CVPR | 3 |
| 2010 | Color Adjacency Modeling for Improved Image and Video SegmentationabstractColor models are often used for representing object appearance for foreground segmentation applications. The relationships between colors can be just as useful for object selection. In this paper, we present a method of modeling color adjacency relationships. By using color adjacency models, the importance of an edge in a given application can be determined and scaled accordingly. We apply our model to foreground segmentation of similar images and video. We show that given one previously-segmented image, we can greatly reduce the error when automatically segmenting other images by using our color adjacency model to weight the likelihood that an edge is part of the desired object boundary. Brian L. Price, Bryan S. Morse, Scott Cohen |
ICPR | 3 |
| 2009 | LIVEcut: Learning-based interactive video segmentation by evaluation of multiple propagated cuesabstractVideo sequences contain many cues that may be used to segment objects in them, such as color, gradient, color adjacency, shape, temporal coherence, camera and object motion, and easily-trackable points. This paper introduces LIVEcut, a novel method for interactively selecting objects in video sequences by extracting and leveraging as much of this information as possible. Using a graph-cut optimization framework, LIVEcut propagates the selection forward frame by frame, allowing the user to correct any mistakes along the way if needed. Enhanced methods of extracting many of the features are provided. In order to use the most accurate information from the various potentially-conflicting features, each feature is automatically weighted locally based on its estimated accuracy using the previous implicitly-validated frame. Feature weights are further updated by learning from the user corrections required in the previous frame. The effectiveness of LIVEcut is shown through timing comparisons to other interactive methods, accuracy comparisons to unsupervised methods, and qualitatively through selections on various video sequences. Brian L. Price, Bryan S. Morse, Scott Cohen |
ICCV | 3 |
| 2005 | Background Estimation as a Labeling ProblemabstractWe present a new background estimation algorithm that constructs the background of an image sequence with moving objects by copying areas from input frames. The background estimation problem is formulated as an optimal labeling problem in which the label at an output pixel is the frame number from which to copy the background color. The costs of assigning labels encourage seamless copying from regions that are stationary over a period of time in such a way that implied motion boundaries occur at intensity edges. This is accomplished without explicitly tracking the moving objects or computing optical flow. Experiments demonstrate that our algorithm is effective in difficult areas where the background is visible for only a small fraction of time, and on inputs with both moving objects that are not always in motion and moving objects with textureless areas Scott Cohen |
ICCV | 1 |