Brian L. Price

dblp:38/5397 · DBLP profile ↗
← Back
84ranked-venue papers
6as first author
22since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 73 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 65 · 5 first-author · 20 since 2021Databases, data management, data science and information retrieval · 6 · 1 since 2021
YearPublicationVenuePosition
2026 Plot'n Polish: Zero-Shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
abstract
Text-to-image diffusion models have demonstrated significant capabilities to generate diverse and detailed visuals in various domains, and story visualization is emerging as a particularly promising application. However, as their use in real-world creative domains increases, the need for providing enhanced control, refinement, and the ability to modify images post-generation in a consistent manner becomes an important challenge. Existing methods often lack the flexibility to apply fine or coarse edits while maintaining visual and narrative consistency across multiple frames, preventing creators from seamlessly crafting and refining their visual stories. To address these challenges, we introduce Plot'n Polish, a zero-shot framework that enables consistent story generation and provides fine-grained control over story visualizations at various levels of detail.
Kiymet Akdemir, Jing Shi 0005, Kushal Kafle, Brian L. Price, Pinar Yanardag Delul
AAAI4
2025 Polarized Color Screen Matting
abstract
This paper considers the long-standing problem of extracting alpha mattes from video using a known background. While various color-based or polarization-based approaches have been studied in past decades, the problem remains ill-posed because the solutions solely rely on either color or polarization. We introduce Polarized Color Screen Matting, a single-shot, per-pixel matting theory for alpha matte and foreground color recovery using both color and polarization cues. Through a theoretical analysis of our diffuse-specular polarimetric compositing equation, we derive practical closed-form matting methods with their solvability conditions. Our theory concludes that an alpha matte can be extracted without manual corrections using off-the-shelf equipment such as an LCD monitor, polarization camera, and unpolarized lights with calibrated color. Experiments on synthetic and real-world datasets verify the validity of our theory and show the capability of our matting methods on real videos with quantitative and qualitative comparisons to color-based and polarization-based matting methods.
Kenji Enomoto, Scott Cohen, Brian L. Price, T. J. Rhodes
CVPR3
2025 Refine-by-Align: Reference-Guided Artifacts Refinement through Semantic Alignment
abstract
Personalized image generation has emerged from the recent advancements in generative models. However, these generated personalized images often suffer from localized artifacts such as incorrect logos, reducing fidelity and fine-grained identity details of the generated results. Furthermore, there is little prior work tackling this problem. To help improve these identity details in the personalized image generation, we introduce a new task: reference-guided artifacts refinement. We present Refine-by-Align, a first-of-its-kind model that employs a diffusion-based framework to address this challenge. Our model consists of two stages: Alignment Stage and Refinement Stage, which share weights of a unified neural network model. Given a generated image, a masked artifact region, and a reference image, the alignment stage identifies and extracts the corresponding regional features in the reference, which are then used by the refinement stage to fix the artifacts. Our model-agnostic pipeline requires no test-time tuning or optimization. It automatically enhances image fidelity and reference identity in the generated image, generalizing well to existing models on various tasks including but not limited to customization, generative compositing, view synthesis, and virtual try-on. Extensive experiments and comparisons demonstrate that our pipeline greatly pushes the boundary of fine details in the image synthesis models.
Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Zhe Lin 0001, Brian L. Price, Scott Cohen, Jianming Zhang 0001, Daniel G. Aliaga
ICLR8
2024 MAGICK: A Large-Scale Captioned Dataset from Matting Generated Images Using Chroma Keying
abstract
We introduce MAGICK, a large-scale dataset of generated objects with high-quality alpha mattes. While image generation methods have produced segmentations, they cannot generate alpha mattes with accurate details in hair, fur, and transparencies. This is likely due to the small size of current alpha matting datasets and the difficulty in obtaining ground-truth alpha. We propose a scalable method for synthesizing images of objects with high-quality alpha that can be used as a ground-truth dataset. A key idea is to generate objects on a single-colored background so chroma keying approaches can be used to extract the alpha. However, this faces several challenges, including that current text-to-image generation methods cannot create images that can be easily chroma keyed and that chroma keying is an underconstrained problem that generally requires manual intervention for high-quality results. We address this using a combination of generation and alpha extraction methods. Using our method, we generate a dataset of 150,000 objects with alpha. We show the utility of our dataset by training an alpha-to-rgb generation method that outperforms baselines. Please see our project website at https://ryanndagreat.github.io/MAGICK/.
Ryan D. Burgert, Brian L. Price, Jason Kuen, Michael S. Ryoo
CVPR2
2024 Putting the Object Back into Video Object Segmentation
abstract
We present Cutie, a video object segmentation (VOS) network with object-level memory reading, which puts the object representation from memory back into the video object segmentation result. Recent works on VOS employ bottom-up pixel-level memory reading which struggles due to matching noise, especially in the presence of distractors, resulting in lower performance in more challenging data. In contrast, Cutie performs top-down object-level memory reading by adapting a small set of object queries. Via those, it interacts with the bottom-up pixel features iteratively with a query-based object transformer (qt, hence Cutie). The object queries act as a high-level summary of the target object, while high-resolution feature maps are retained for accurate segmentation. Together with foreground-background masked attention, Cutie cleanly separates the semantics of the foreground object from the background. On the challenging MOSE dataset, Cutie improves by 8.7$\mathcal{J}\& \mathcal{F}$over XMem with a similar running time and improves by 4.2$\mathcal{J}\&\mathcal{F}$over DeAOT while being three times faster. Code is available at: hkchengrex.github.io/Cutie.
Ho Kei Cheng, Seoung Wug Oh, Brian L. Price, Joon-Young Lee, Alexander G. Schwing
CVPR3
2024 Polar Matte: Fully Computational Ground-Truth-Quality Alpha Matte Extraction for Images and Video using Polarized Screen Matting
Kenji Enomoto, T. J. Rhodes, Brian L. Price, Gavin Miller
CVPR3
2024 IMPRINT: Generative Object Compositing by Learning Identity-Preserving Representation
abstract
Generative object compositing emerges as a promising new avenue for compositional image editing. However, the requirement of object identity preservation poses a significant challenge, limiting practical usage of most existing methods. In response, this paper introduces IMPRINT, a novel diffusion-based generative model trained with a two-stage learning framework that decouples learning of identity preservation from that of compositing. The first stage is targeted for context-agnostic, identity-preserving pretraining of the object encoder, enabling the encoder to learn an embedding that is both view-invariant and conducive to enhanced detail preservation. The subsequent stage leverages this representation to learn seamless harmonization of the object composited to the background. In addition, IMPRINT incorporates a shape-guidance mechanism offering user-directed control over the compositing process. Extensive experiments demonstrate that IMPRINT significantly outperforms existing methods and various baselines on identity preservation and composition quality. Project page: https://song630.github.io/IMPRINT-Project-Page/
Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, He Zhang 0004, Wei Xiong 0008, Daniel G. Aliaga
CVPR5
2024 SPIN: Hierarchical Segmentation with Subpart Granularity in Natural Images
Josh Myers-Dean, Jarek Reynolds, Brian L. Price, Danna Gurari
ECCV (24)3
2024 SegGen: Supercharging Segmentation Models with Text2Mask and Mask2Img Synthesis
Hanrong Ye, Jason Kuen, Qing Liu 0017, Zhe Lin 0001, Brian L. Price, Dan Xu 0002
ECCV (8)5
2024 Uncertainty-aware Fine-tuning of Segmentation Foundation Models
abstract
The Segment Anything Model (SAM) is a large-scale foundation model that has revolutionized segmentation methodology. Despite its impressive generalization ability, the segmentation accuracy of SAM on images with intricate structures is often unsatisfactory. Recent works have proposed lightweight fine-tuning using high-quality annotated data to improve accuracy on such images. However, here we provide extensive empirical evidence that this strategy leads to forgetting how to "segment anything": these models lose the original generalization abilities of SAM, in the sense that they perform worse for segmentation tasks not represented in the annotated fine-tuning set. To improve performance without forgetting, we introduce a novel framework that combines high-quality annotated data with a large unlabeled dataset. The framework relies on two methodological innovations. First, we quantify the uncertainty in the SAM pseudo labels associated with the unlabeled data and leverage it to perform uncertainty-aware fine-tuning. Second, we encode the type of segmentation task associated with each training example using a $\textit{task prompt}$ to reduce ambiguity. We evaluated the proposed Segmentation with Uncertainty Model (SUM) on a diverse test set consisting of 14 public benchmarks, where it achieves state-of-the-art results. Notably, our method consistently surpasses SAM by 3-6 points in mean IoU and 4-7 in mean boundary IoU across point-prompt interactive segmentation rounds. Code is available at https://github.com/Kangningthu/SUM
Kangning Liu, Brian L. Price, Jason Kuen, Zijun Wei, Luis Figueroa, Krzysztof J. Geras, Carlos Fernandez-Granda
NeurIPS2
2024 Interactive Segmentation for Diverse Gesture Types Without Context
abstract
Interactive segmentation entails a human marking an image to guide how a model either creates or edits a segmentation. Our work addresses limitations of existing methods: they either only support one gesture type for marking an image (e.g., either clicks or scribbles) or require knowledge of the gesture type being employed, and require specifying whether marked regions should be included versus excluded in the final segmentation. We instead propose a simplified interactive segmentation task where a user only must mark an image, where the input can be of any gesture type without specifying the gesture type. We support this new task by introducing the first interactive segmentation dataset with multiple gesture types as well as a new evaluation metric capable of holistically evaluating interactive segmentation algorithms. We then analyze numerous interactive segmentation algorithms, including ones adapted for our novel task. While we observe promising performance overall, we also highlight areas for future improvement. To facilitate further extensions of this work, we publicly share our new dataset at https://github.com/joshmyersdean/dig.
Josh Myers-Dean, Brian L. Price, Wilson Chan, Danna Gurari
WACV3
2023 GamutMLP: A Lightweight MLP for Color Loss Recovery
abstract
Cameras and image-editing software often process images in the wide-gamut ProPhoto color space, encompassing 90% of all visible colors. However, when images are encoded for sharing, this color-rich representation is transformed and clipped to fit within the small-gamut standard RGB (sRGB) color space, representing only 30% of visible colors. Recovering the lost color information is challenging due to the clipping procedure. Inspired by neural implicit representations for 2D images, we propose a method that optimizes a lightweight multi-layer-perceptron (MLP) model during the gamut reduction step to predict the clipped values. GamutMLP takes approximately 2 seconds to optimize and requires only 23 KB of storage. The small memory footprint allows our GamutMLP model to be saved as metadata in the sRGB image—the model can be extracted when needed to restore wide-gamut color values. We demonstrate the effectiveness of our approach for color recovery and compare it with alternative strategies, including pre-trained DNN-based gamut expansion networks and other implicit neural representation methods. As part of this effort, we introduce a new color gamut dataset of 2200 wide-gamut/small-gamut images for training and testing.
Hoang Minh Le 0001, Brian L. Price, Scott Cohen, Michael S. Brown
CVPR2
2023 Towards Open-World Segmentation of Parts
abstract
Segmenting object parts such as cup handles and animal bodies is important in many real-world applications but requires more annotation effort. The largest dataset nowadays contains merely two hundred object categories, implying the difficulty to scale up part segmentation to an unconstrained setting. To address this, we propose to explore a seemingly simplified but empirically useful and scalable task, class-agnostic part segmentation. In this problem, we disregard the part class labels in training and instead treat all of them as a single part class. We argue and demonstrate that models trained without part classes can better localize parts and segment them on objects unseen in training. We then present two further improvements. First, we propose to make the model object-aware, leveraging the fact that parts are “compositions”, whose extents are bounded by the corresponding objects and whose appearances are by nature not independent but bundled. Second, we introduce a novel approach to improve part segmentation on unseen objects, inspired by an interesting finding - for unseen objects, the pixel-wise features extracted by the model often reveal high-quality part segments. To this end, we propose a novel self-supervised procedure that iterates between pixel clustering and supervised contrastive learning that pulls pixels closer or pushes them away. Via extensive experiments on PartImageNet and Pascal-Part, we show notable and consistent gains by our approach, essentially a critical step towards open-world part segmentation.
Tai-Yu Pan, Qing Liu 0017, Wei-Lun Chao, Brian L. Price
CVPR4
2023 ObjectStitch: Object Compositing with Diffusion Model
abstract
Object compositing based on 2D images is a challenging problem since it typically involves multiple processing stages such as color harmonization, geometry correction and shadow generation to generate realistic results. Furthermore, annotating training data pairs for compositing requires substantial manual effort from professionals, and is hardly scalable. Thus, with the recent advances in generative models, in this work, we propose a selfsupervised framework for object compositing by leveraging the power of conditional diffusion models. Our framework can hollistically address the object compositing task in a unified model, transforming the viewpoint, geometry, color and shadow of the generated object while requiring no manual labeling. To preserve the input object's characteristics, we introduce a content adaptor that helps to maintain categori-cal semantics and object appearance. A data augmentation method is further adopted to improve the fidelity of the generator. Our method outperforms relevant baselines in both realism and faithfulness of the synthesized result images in a user study on various real-world images.
Zhe Lin 0001, Scott Cohen, Brian L. Price, Jianming Zhang 0001, Soo Ye Kim, Daniel G. Aliaga
CVPR5
2023 Tracking Anything with Decoupled Video Segmentation
abstract
Training data for video segmentation are expensive to annotate. This impedes extensions of end-to-end algorithms to new video segmentation tasks, especially in large-vocabulary settings. To ‘track anything’ without training on video data for every individual task, we develop a decoupled video segmentation approach (DEVA), composed of task-specific image-level segmentation and class/task-agnostic bi-directional temporal propagation. Due to this design, we only need an image-level model for the target task (which is cheaper to train) and a universal temporal propagation model which is trained once and generalizes across tasks. To effectively combine these two modules, we use bi-directional propagation for (semi-)online fusion of segmentation hypotheses from different frames to generate a coherent segmentation. We show that this decoupled formulation compares favorably to end-to-end approaches in several data-scarce tasks including large-vocabulary video panoptic segmentation, open-world video segmentation, referring video segmentation, and unsupervised video object segmentation. Code is available at: hkchengrex.github.io/Tracking-Anything-with-DEVA.
Ho Kei Cheng, Seoung Wug Oh, Brian L. Price, Alexander G. Schwing, Joon-Young Lee
ICCV3
2022 Boosting Robustness of Image Matting with Context Assembling and Strong Data Augmentation
abstract
Deep image matting methods have achieved increasingly better results on benchmarks (e.g., Composition-1k/alphamatting.com). However, the robustness, including robustness to trimaps and generalization to images from different domains, is still underexplored. Although some works propose to either refine the trimaps or adapt the algorithms to real-world images via extra data augmentation, none of them has taken both into consideration, not to mention the significant performance deterioration on benchmarks while using those data augmentation. To fill this gap, we propose an image matting method which achieves higher robustness (RMat) via multilevel context assembling and strong data augmentation targeting matting. Specifically, we first build a strong matting framework by modeling ample global information with transformer blocks in the encoder, and focusing on details in combination with convolution layers as well as a low-level feature assembling attention block in the decoder. Then, based on this strong baseline, we analyze current data augmentation and explore simple but effective strong data augmentation to boost the baseline model and contribute a more generalizable matting method. Compared with previous methods, the proposed method not only achieves state-of-the-art results on the Composition-1k benchmark (11 % improvement on SAD and 27% improvement on Grad) with smaller model size, but also shows more robust generalization results on other benchmarks, on real-world images, and also on varying coarse-to-fine trimaps with our extensive experiments.11This work was in part done when YD was an intern at Adobe and CS was with The University of Adelaide. CS is the corresponding author. Project page: https://dongdong93.github.io/RMat/.
Yutong Dai 0001, Brian L. Price, He Zhang 0004, Chunhua Shen
CVPR2
2022 Generalizing Interactive Backpropagating Refinement for Dense Prediction Networks
abstract
As deep neural networks become the state-of-the-art approach in the field of computer vision for dense prediction tasks, many methods have been developed for automatic estimation of the target outputs given the visual inputs. Although the estimation accuracy of the proposed automatic methods continues to improve, interactive refinement is oftentimes necessary for further correction. Recently, feature backpropagating refinement scheme [25] (f-BRS) has been proposed for the task of interactive segmentation, which enables efficient optimization of a small set of auxiliary variables inserted into the pretrained network to produce object segmentation that better aligns with user inputs. However, the proposed auxiliary variables only contain channel-wise scale and bias, limiting the optimization to global refinement only. In this work, in order to generalize backpropagating refinement for a wide range of dense prediction tasks, we introduce a set of G-BRS (Generalized Backpropagating Refinement Scheme) layers that enable both global and localized refinement for the following tasks: interactive segmentation, semantic segmentation, image matting and monocular depth estimation. Experiments on SBD, Cityscapes, Mapillary Vista, Composition-1k and NYU-Depth-V2 show that our method can successfully generalize and significantly improve performance of existing pretrained state-of-the-art models with only a few clicks.
Fanqing Lin, Brian L. Price, Tony R. Martinez
CVPR2
2022 One-Trimap Video Matting
Hongje Seong, Seoung Wug Oh, Brian L. Price, Euntai Kim, Joon-Young Lee
ECCV (29)3
2022 Measuring Human Perception to Improve Handwritten Document Transcription
abstract
In this paper, we consider how to incorporate psychophysical measurements of human visual perception into the loss function of a deep neural network being trained for a recognition task, under the assumption that such information can reduce errors. As a case study to assess the viability of this approach, we look at the problem of handwritten document transcription. While good progress has been made towards automatically transcribing modern handwriting, significant challenges remain in transcribing historical documents. Here we describe a general enhancement strategy, underpinned by the new loss formulation, which can be applied to the training regime of any deep learning-based document transcription system. Through experimentation, reliable performance improvement is demonstrated for the standard IAM and RIMES datasets for three different network architectures. Further, we go on to show feasibility for our approach on a new dataset of digitized Latin manuscripts, originally produced by scribes in the Cloister of St. Gall in the the 9th century.
Samuel Grieggs, Bingyu Shen 0001, Greta Rauch, David Chiang 0001, Brian L. Price, Walter J. Scheirer
IEEE Trans. Pattern Anal. Mach. Intell.7
2021 Rethinking Text Segmentation: A Novel Dataset and a Text-Specific Refinement Approach
abstract
Text segmentation is a prerequisite in many real-world text-related tasks, e.g., text style transfer, and scene text removal. However, facing the lack of high-quality datasets and dedicated investigations, this critical prerequisite has been left as an assumption in many works, and has been largely overlooked by current research. To bridge this gap, we proposed TextSeg, a large-scale fine-annotated text dataset with six types of annotations: word- and character-wise bounding polygons, masks, and transcriptions. We also introduce Text Refinement Network (TexRNet), a novel text segmentation approach that adapts to the unique properties of text, e.g. non-convex boundary, diverse texture, etc., which often impose burdens on traditional segmentation models. In our TexRNet, we propose text-specific network designs to address such challenges, including key features pooling and attention-based similarity checking. We also introduce trimap and discriminator losses that show significant improvement in text segmentation. Extensive experiments are carried out on both our TextSeg dataset and other existing datasets. We demonstrate that TexRNet consistently improves text segmentation performance by nearly 2% compared to other state-of-the-art segmentation methods. Our dataset and code can be found at https://github.com/SHI-Labs/Rethinking-TextSegmentation.
Xingqian Xu, Brian L. Price, Zhonghao Wang 0001, Humphrey Shi
CVPR4
2021 Visual FUDGE: Form Understanding via Dynamic Graph Editing
Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington
ICDAR (1)3
2021 Deep Interactive Thin Object Selection
abstract
Existing deep learning based interactive segmentation methods have achieved remarkable performance with only a few user clicks, e.g. DEXTR [32] attaining 91.5% IoU on PASCAL VOC with only four extreme clicks. However, we observe even the state-of-the-art methods would often struggle in cases of objects to be segmented with elongated thin structures (e.g. bug legs and bicycle spokes). We investigate such failures, and find the critical reasons behind are two-fold: 1) lack of appropriate training dataset; and 2) extremely imbalanced distribution w.r.t. number of pixels belonging to thin and non-thin regions. Targeted at these challenges, we collect a large-scale dataset specifically for segmentation of thin elongated objects, named ThinObject-5K. Also, we present a novel integrative thin object segmentation network consisting of three streams. Among them, the high-resolution edge stream aims at preserving fine-grained details including elongated thin parts; the fixed-resolution context stream focuses on capturing semantic contexts. The two streams' outputs are then amalgamated in the fusion stream to complement each other for help producing a refined segmentation output with sharper predictions around thin parts. Extensive experimental results well demonstrate the effectiveness of our proposed solution on segmenting thin objects, surpassing the baseline by ~ 30% IoUthindespite using only four clicks. Codes and dataset are available at https://github.com/liewjunhao/thin-object-selection.
Jun Hao Liew, Scott Cohen, Brian L. Price, Long Mai, Jiashi Feng
WACV3
2020 Text and Style Conditioned GAN for the Generation of Offline-Handwriting Lines
Brian L. Davis, Bryan S. Morse, Brian L. Price, Chris Tensmeyer, Curtis Wigington, Rajiv Jain
BMVC3
2020 Deepstrip: High-Resolution Boundary Refinement
abstract
In this paper, we target refining the boundaries in high resolution images given low resolution masks. For memory and computation efficiency, we propose to convert the regions of interest into strip images and compute a boundary prediction in the strip domain. To detect the target boundary, we present a framework with two prediction layers. First, all potential boundaries are predicted as an initial prediction and then a selection layer is used to pick the target boundary and smooth the result. To encourage accurate prediction, a loss which measures the boundary distance in strip domain is introduced. In addition, we enforce a matching consistency and C0 continuity regularization to the network to reduce false alarms. Extensive experiments on both public and a newly created high resolution dataset strongly validate our approach.
Peng Zhou 0009, Brian L. Price, Scott Cohen, Gregg Wilensky, Larry Davis 0001
CVPR2
2020 PhraseClick: Toward Achieving Flexible Interactive Segmentation by Phrase and Click
Henghui Ding, Scott Cohen, Brian L. Price, Xudong Jiang 0001
ECCV (3)3
2020 Interactive Training And Architecture For Deep Object Selection
abstract
Interactive object cutout tools are the cornerstone of the image editing workflow. Algorithms that can reduce the number of interactions are clearly valuable. Recent deep-learning based interactive segmentation algorithms are capable of rough binary selections with a handful of clicks, yet, they tend to plateau once this rough selection has been reached. In this work, we interpret this plateau as an inability of the algorithm to precisely leverage each user interaction.We introduce a novel interactive architecture and a training scheme that are both tailored to better exploit the user input at higher numbers of clicks. Comprehensive experiments support our approach, and our network achieves state of the art performance.
Marco Forte, Brian L. Price, Scott Cohen, Ning Xu 0007, François Pitié
ICME2
2020 Answering Questions about Data Visualizations using Efficient Bimodal Fusion
abstract
Chart question answering (CQA) is a newly proposed visual question answering (VQA) task where an algorithm must answer questions about data visualizations, e.g. bar charts, pie charts, and line graphs. CQA requires capabilities that natural-image VQA algorithms lack: fine-grained measurements, optical character recognition, and handling out-of-vocabulary words in both questions and answers. Without modifications, state-of-the-art VQA algorithms perform poorly on this task. Here, we propose a novel CQA algorithm called parallel recurrent fusion of image and language (PReFIL). PReFIL first learns bimodal embeddings by fusing question and image features and then intelligently aggregates these learned embeddings to answer the given question. Despite its simplicity, PReFIL greatly surpasses state-of-the art systems and human baselines on both the FigureQA and DVQA datasets. Additionally, we demonstrate that PReFIL can be used to reconstruct tables by asking a series of questions about a chart.
Kushal Kafle, Robik Shrestha, Brian L. Price, Scott Cohen, Christopher Kanan
WACV3
2020 RGB2AO: Ambient Occlusion Generation from RGB Images
abstract
Abstract We present RGB2AO, a novel task to generate ambient occlusion (AO) from a single RGB image instead of screen space buffers such as depth and normal. RGB2AO produces a new image filter that creates a non‐directional shading effect that darkens enclosed and sheltered areas. RGB2AO aims to enhance two 2D image editing applications: image composition and geometry‐aware contrast enhancement. We first collect a synthetic dataset consisting of pairs of RGB images and AO maps. Subsequently, we propose a model for RGB2AO by supervised learning of a convolutional neural network (CNN), considering 3D geometry of the input image. Experimental results quantitatively and qualitatively demonstrate the effectiveness of our model.
Naoto Inoue, Daichi Ito, Yannick Hold-Geoffroy, Long Mai, Brian L. Price, Toshihiko Yamasaki
Comput. Graph. Forum5
2019 When Color Constancy Goes Wrong: Correcting Improperly White-Balanced Images
abstract
This paper focuses on correcting a camera image that has been improperly white-balanced. This situation occurs when a camera's auto white balance fails or when the wrong manual white-balance setting is used. Even after decades of computational color constancy research, there are no effective solutions to this problem. The challenge lies not in identifying what the correct white balance should have been, but in the fact that the in-camera white-balance procedure is followed by several camera-specific nonlinear color manipulations that make it challenging to correct the image's colors in post-processing. This paper introduces the first method to explicitly address this problem. Our method is enabled by a dataset of over 65,000 pairs of incorrectly white-balanced images and their corresponding correctly white-balanced images. Using this dataset, we introduce a k-nearest neighbor strategy that is able to compute a nonlinear color mapping function to correct the image's colors. We show our method is highly effective and generalizes well to camera models not in the training set.
Mahmoud Afifi, Brian L. Price, Scott Cohen, Michael S. Brown
CVPR2
2019 MultiSeg: Semantically Meaningful, Scale-Diverse Segmentations From Minimal User Input
abstract
Existing deep learning-based interactive image segmentation approaches typically assume the target-of-interest is always a single object and fail to account for the potential diversity in user expectations, thus requiring excessive user input when it comes to segmenting an object part or a group of objects instead. Motivated by the observation that the object part, full object, and a collection of objects essentially differ in size, we propose a new concept called scale-diversity, which characterizes the spectrum of segmentations w.r.t. different scales. To address this, we present MultiSeg, a scale-diverse interactive image segmentation network that incorporates a set of two-dimensional scale priors into the model to generate a set of scale-varying proposals that conform to the user input. We explicitly encourage segmentation diversity during training by synthesizing diverse training samples for a given image. As a result, our method allows the user to quickly locate the closest segmentation target for further refinement if necessary. Despite its simplicity, experimental results demonstrate that our proposed model is capable of quickly producing diverse yet plausible segmentation outputs, reducing the user interaction required, especially in cases where many types of segmentations (object parts or groups) are expected.
Jun Hao Liew, Scott Cohen, Brian L. Price, Long Mai, Sim Heng Ong, Jiashi Feng
ICCV3
2019 Unconstrained Foreground Object Search
abstract
Many people search for foreground objects to use when editing images. While existing methods can retrieve candidates to aid in this, they are constrained to returning objects that belong to a pre-specified semantic class. We instead propose a novel problem of unconstrained foreground object (UFO) search and introduce a solution that supports efficient search by encoding the background image in the same latent space as the candidate foreground objects. A key contribution of our work is a cost-free, scalable approach for creating a large-scale training dataset with a variety of foreground objects of differing semantic categories per image location. Quantitative and human-perception experiments with two diverse datasets demonstrate the advantage of our UFO search solution over related baselines.
Brian L. Price, Scott Cohen, Danna Gurari
ICCV2
2019 Deep Visual Template-Free Form Parsing
abstract
The following topics are dealt with: learning (artificial intelligence); document image processing; feature extraction; text analysis; convolutional neural nets; image segmentation; handwritten character recognition; image classification; text detection; optical character recognition.
Brian L. Davis, Bryan S. Morse, Scott Cohen, Brian L. Price, Chris Tensmeyer
ICDAR4
2019 Deep Splitting and Merging for Table Structure Decomposition
abstract
Given the large variety and complexity of tables, table structure extraction is a challenging task in automated document analysis systems. We present a pair of novel deep learning models (Split and Merge models) that given an input image, 1) predicts the basic table grid pattern and 2) predicts which grid elements should be merged to recover cells that span multiple rows or columns. We propose projection pooling as a novel component of the Split model and grid pooling as a novel part of the Merge model. While most Fully Convolutional Networks rely on local evidence, these unique pooling regions allow our models to take advantage of the global table structure. We achieve state-of-the-art performance on the public ICDAR 2013 Table Competition dataset of PDF documents. On a much larger private dataset which we used to train the models, we significantly outperform both a state-ofthe-art deep model and a major commercial software system.
Chris Tensmeyer, Vlad I. Morariu, Brian L. Price, Scott Cohen, Tony R. Martinez
ICDAR3
2019 Multi-label Connectionist Temporal Classification
abstract
The Connectionist Temporal Classification (CTC) loss function [1] enables end-to-end training of a neural network for sequence-to-sequence tasks without the need for prior alignments between the input and output. CTC is traditionally used for training sequential, single-label problems; each element in the sequence has only one class. In this work, we show that CTC is not suitable for multi-label tasks and we present a novel Multi-label Connectionist Temporal Classification (MCTC) loss function for multi-label, sequence-to-sequence classification. Multi-label classes can represent meaningful attributes of a single element; for example, in Optical Music Recognition (OMR), a music note can have separate duration and pitch attributes. Our approach achieves state-of-the-art results on Joint Handwritten Text Recognition and Name Entity Recognition, Asian Character Recognition, and OMR.
Curtis Wigington, Brian L. Price, Scott Cohen
ICDAR2
2019 Guided Image Inpainting: Replacing an Image Region by Pulling Content From Another Image
abstract
Deep generative models have shown success in automatically synthesizing missing image regions using surrounding context. However, users cannot directly decide what content to synthesize with such approaches.We propose an end-to-end network for image inpainting that uses a different image to guide the synthesis of new content to fill the hole. A key challenge addressed by our approach is synthesizing new content in regions where the guidance image and the context of the original image are inconsistent. We conduct four studies that demonstrate our method yields more realistic image inpainting results over seven baselines.
Brian L. Price, Scott Cohen, Danna Gurari
WACV2
2019 Learning to Trace: Expressive Line Drawing Generation from Photographs
abstract
Abstract In this paper, we present a new computational method for automatically tracing high‐resolution photographs to create expressive line drawings. We define expressive lines as those that convey important edges, shape contours, and large‐scale texture lines that are necessary to accurately depict the overall structure of objects (similar to those found in technical drawings) while still being sparse and artistically pleasing. Given a photograph, our algorithm extracts expressive edges and creates a clean line drawing using a convolutional neural network (CNN). We employ an end‐to‐end trainable fully‐convolutional CNN to learn the model in a data‐driven manner. The model consists of two networks to cope with two sub‐tasks; extracting coarse lines and refining them to be more clean and expressive. To build a model that is optimal for each domain, we construct two new datasets for face/body and manga background. The experimental results qualitatively and quantitatively demonstrate the effectiveness of our model. We further illustrate two practical applications.
Naoto Inoue, Daichi Ito, Ning Xu 0007, Brian L. Price, Toshihiko Yamasaki
Comput. Graph. Forum5
2018 Disentangling Structure and Aesthetics for Style-Aware Image Completion
abstract
Content-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two-fold: (1) we explicitly disentangle image structure and style during patch search and selection to ensure a visually consistent look and feel within the target image. (2) we perform adaptive stylization of patches to conform the aesthetics of selected patches to the target image, so harmonizing the integration of selected patches into the final composition. We show that explicit consideration of visual style during in-painting delivers excellent qualitative and quantitative results across the varied image styles and content, over the Places2 scene photographic dataset and a challenging new in-painting dataset of artwork derived from BAM!
Andrew Gilbert, John P. Collomosse, Hailin Jin, Brian L. Price
CVPR4
2018 DVQA: Understanding Data Visualizations via Question Answering
abstract
Bar charts are an effective way to convey numeric information, but today's algorithms cannot parse them. Existing methods fail when faced with even minor variations in appearance. Here, we present DVQA, a dataset that tests many aspects of bar chart understanding in a question answering framework. Unlike visual question answering (VQA), DVQA requires processing words and answers that are unique to a particular bar chart. State-of-the-art VQA algorithms perform poorly on DVQA, and we propose two strong baselines that perform considerably better. Our work will enable algorithms to automatically extract numeric and semantic information from vast quantities of bar charts found in scientific publications, Internet articles, business reports, and many other areas.
Kushal Kafle, Brian L. Price, Scott Cohen, Christopher Kanan
CVPR2
2018 Discriminability Objective for Training Descriptive Captions
abstract
One property that remains lacking in image captions generated by contemporary methods is discriminability: being able to tell two images apart given the caption for one of them. We propose a way to improve this aspect of caption generation. By incorporating into the captioning training objective a loss component directly related to ability (by a machine) to disambiguate image/caption matches, we obtain systems that produce much more discriminative caption, according to human evaluation. Remarkably, our approach leads to improvement in other aspects of generated captions, reflected by a battery of standard scores such as BLEU, SPICE etc. Our approach is modular and can be applied to a variety of model/loss combinations commonly proposed for image captioning.
Ruotian Luo, Brian L. Price, Scott Cohen, Gregory Shakhnarovich
CVPR2
2018 Interactive Boundary Prediction for Object Selection
Hoang Le, Long Mai, Brian L. Price, Scott Cohen, Hailin Jin, Feng Liu 0015
ECCV (14)3
2018 Start, Follow, Read: End-to-End Full-Page Handwriting Recognition
Curtis Wigington, Chris Tensmeyer, Brian L. Davis, Bill Barrett, Brian L. Price, Scott Cohen
ECCV (6)5
2018 YouTube-VOS: Sequence-to-Sequence Video Object Segmentation
Ning Xu 0007, Yuchen Fan 0001, Jianchao Yang, Dingcheng Yue, Brian L. Price, Scott Cohen, Thomas S. Huang
ECCV (5)7
2018 Compositing-Aware Image Search
Hengshuang Zhao, Xiaohui Shen, Zhe Lin 0001, Kalyan Sunkavalli, Brian L. Price, Jiaya Jia
ECCV (3)5
2017 Sherlock: Scalable Fact Learning in Images
abstract
We study scalable and uniform understanding of facts in images. Existing visual recognition systems are typically modeled differently for each fact type such as objects, actions, and interactions. We propose a setting where all these facts can be modeled simultaneously with a capacity to understand an unbounded number of facts in a structured way. The training data comes as structured facts in images, including (1) objects (e.g., ), (2) attributes (e.g., ), (3) actions (e.g., ), and (4) interactions (e.g., ). Each fact has a semantic language view (e.g., < boy, playing>) and a visual view (an image with this fact). We show that learning visual facts in a structured way enables not only a uniform but also generalizable visual understanding. We propose and investigate recent and strong approaches from the multiview learning literature and also introduce two learning representation models as potential baselines. We applied the investigated methods on several datasets that we augmented with structured facts and a large scale dataset of more than 202,000 facts and 814,000 images. Our experiments show the advantage of relating facts by the structure by the proposed models compared to the designed baselines on bidirectional fact retrieval.
Scott Cohen, Walter Chang, Brian L. Price, Ahmed M. Elgammal
AAAI4
2017 Deep GrabCut for Object Selection
Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang
BMVC2
2017 Forecasting Human Dynamics from Static Images
abstract
This paper presents the first study on forecasting human dynamics from static images. The problem is to input a single RGB image and generate a sequence of upcoming human body poses in 3D. To address the problem, we propose the 3D Pose Forecasting Network (3D-PFNet). Our 3D-PFNet integrates recent advances on single-image human pose estimation and sequence prediction, and converts the 2D predictions into 3D space. We train our 3D-PFNet using a three-step training strategy to leverage a diverse source of training data, including image and video based human pose datasets and 3D motion capture (MoCap) data. We demonstrate competitive performance of our 3D-PFNet on 2D pose forecasting and 3D structure recovery through quantitative and qualitative results.
Yu-Wei Chao, Jimei Yang, Brian L. Price, Scott Cohen, Jia Deng 0001
CVPR3
2017 Depth from Defocus in the Wild
abstract
We consider the problem of two-frame depth from defocus in conditions unsuitable for existing methods yet typical of everyday photography: a non-stationary scene, a handheld cellphone camera, a small aperture, and sparse scene texture. The key idea of our approach is to combine local estimation of depth and flow in very small patches with a global analysis of image content-3D surfaces, deformations, figure-ground relations, textures. To enable local estimation we (1) derive novel defocus-equalization filters that induce brightness constancy across frames and (2) impose a tight upper bound on defocus blur-just three pixels in radius-by appropriately refocusing the camera for the second input frame. For global analysis we use a novel splinebased scene representation that can propagate depth and flow across large irregularly-shaped regions. Our experiments show that this combination preserves sharp boundaries and yields good depth and flow maps in the face of significant noise, non-rigidity, and data sparsity.
Huixuan Tang, Scott Cohen, Brian L. Price, Stephen Schiller, Kiriakos N. Kutulakos
CVPR3
2017 Deep Image Matting
abstract
Image matting is a fundamental computer vision problem and has many applications. Previous algorithms have poor performance when an image has similar foreground and background colors or complicated textures. The main reasons are prior methods 1) only use low-level features and 2) lack high-level context. In this paper, we propose a novel deep learning based algorithm that can tackle both these problems. Our deep model has two parts. The first part is a deep convolutional encoder-decoder network that takes an image and the corresponding trimap as inputs and predict the alpha matte of the image. The second part is a small convolutional network that refines the alpha matte predictions of the first network to have more accurate alpha values and sharper edges. In addition, we also create a large-scale image matting dataset including 49300 training images and 1000 testing images. We evaluate our algorithm on the image matting benchmark, our testing set, and a wide variety of real images. Experimental results clearly demonstrate the superiority of our algorithm over previous methods.
Ning Xu 0007, Brian L. Price, Scott Cohen, Thomas S. Huang
CVPR2
2017 Multi-Scale Multi-Task FCN for Semantic Page Segmentation and Table Detection
abstract
Page segmentation and table detection play an important role in understanding the structure of documents. We present a page segmentation algorithm that incorporates state-of-the-art deep learning methods for segmenting three types of document elements: text blocks, tables, and figures. We propose a multi-scale, multi-task fully convolutional neural network (FCN) for the tasks of semantic page segmentation and element contour detection. The semantic segmentation network accurately predicts the probability at each pixel of the three element classes. The contour detection network accurately predicts instance level "edges" around each element occurrence. We propose a conditional random field (CRF) that uses features output from the semantic segmentation and contour networks to improve upon the semantic segmentation network output. Given the semantic segmentation output, we also extract individual table instances from the page using some heuristic rules and a verification network to remove false positives. We show that although we only consider a page image as input, we produce comparable results with other methods that relies on PDF file information and heuristics and hand crafted features tailored to specific types of documents. Our approach learns the representative features for page segmentation from real and synthetic training data. %, and produces good results on real documents. The learning-based property makes it a more general method than existing methods in terms of document types and element appearances. For example, our method reliably detects sparsely lined tables which are hard for rule-based or heuristic methods.
Dafang He, Scott Cohen, Brian L. Price, Daniel Kifer, C. Lee Giles
ICDAR3
2017 Data Augmentation for Recognition of Handwritten Words and Lines Using a CNN-LSTM Network
abstract
We introduce two data augmentation and normalization techniques, which, used with a CNN-LSTM, significantly reduce Word Error Rate (WER) and Character Error Rate (CER) beyond best-reported results on handwriting recognition tasks. (1) We apply a novel profile normalization technique to both word and line images. (2) We augment existing text images using random perturbations on a regular grid. We apply our normalization and augmentation to both training and test images. Our approach achieves low WER and CER over hundreds of authors, multiple languages and a variety of collections written centuries apart. Image augmentation in this manner achieves state-of-the-art recognition accuracy on several popular handwritten word benchmarks.
Curtis Wigington, Seth Stewart, Brian L. Davis, Bill Barrett, Brian L. Price, Scott Cohen
ICDAR5
2017 Group-Theme Recoloring for Multi-Image Color Consistency
abstract
Abstract Modifying the colors of an image is a fundamental editing task with a wide range of methods available. Manipulating multiple images to share similar colors is more challenging, with limited tools available. Methods such as color transfer are effective in making an image share similar colors with a target image; however, color transfer is not suitable for modifying multiple images. Approaches for color consistency for photo collections give good results when the photo collection contains similar scene content, but are not applicable for general input images. To address these gaps, we propose an application framework for achieving color consistency for multi‐image input. Our framework derives a group color theme from the input images′ individual color palettes and uses this group color theme to recolor the image collection. This group‐theme recoloring provides an effective way to ensure color consistency among multiple images and naturally lends itself to the inclusion of an additional external color theme. We detail our group‐theme recoloring approach and demonstrate its effectiveness on a number of examples.
Nguyen Ho Man Rang, Brian L. Price, Scott Cohen, Michael S. Brown
Comput. Graph. Forum2
2017 Salient Object Subitizing
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
Int. J. Comput. Vis.8
2016 Two Illuminant Estimation and User Correction Preference
abstract
This paper examines the problem of white-balance correction when a scene contains two illuminations. This is a two step process: 1) estimate the two illuminants, and 2) correct the image. Existing methods attempt to estimate a spatially varying illumination map, however, results are error prone and the resulting illumination maps are too lowresolution to be used for proper spatially varying whitebalance correction. In addition, the spatially varying nature of these methods make them computationally intensive. We show that this problem can be effectively addressed by not attempting to obtain a spatially varying illumination map, but instead by performing illumination estimation on large sub-regions of the image. Our approach is able to detect when distinct illuminations are present in the image and accurately measure these illuminants. Since our proposed strategy is not suitable for spatially varying image correction, a user study is performed to see if there is a preference for how the image should be corrected when two illuminants are present, but only a global correction can be applied. The user study shows that when the illuminations are distinct, there is a preference for the outdoor illumination to be corrected resulting in warmer final result. We use these collective findings to demonstrate an effective two illuminant estimation scheme that produces corrected images that users prefer.
Dongliang Cheng, Abdelrahman Kamel, Brian L. Price, Scott Cohen, Michael S. Brown
CVPR3
2016 Interactive Segmentation on RGBD Images via Cue Selection
abstract
Interactive image segmentation is an important problem in computer vision with many applications including image editing, object recognition and image retrieval. Most existing interactive segmentation methods only operate on color images. Until recently, very few works have been proposed to leverage depth information from low-cost sensors to improve interactive segmentation. While these methods achieve better results than color-based methods, they are still limited in either using depth as an additional color channel or simply combining depth with color in a linear way. We propose a novel interactive segmentation algorithm which can incorporate multiple feature cues like color, depth, and normals in an unified graph cut framework to leverage these cues more effectively. A key contribution of our method is that it automatically selects a single cue to be used at each pixel, based on the intuition that only one cue is necessary to determine the segmentation label locally. This is achieved by optimizing over both segmentation labels and cue labels, using terms designed to decide where both the segmentation and label cues should change. Our algorithm thus produces not only the segmentation mask but also a cue label map that indicates where each cue contributes to the final result. Extensive experiments on five large scale RGBD datasets show that our proposed algorithm performs significantly better than both other color-based and RGBD based algorithms in reducing the amount of user inputs as well as increasing segmentation accuracy.
Brian L. Price, Scott Cohen, Shih-Fu Chang
CVPR2
2016 Deep Interactive Object Selection
abstract
Interactive object selection is a very important research problem and has many applications. Previous algorithms require substantial user interactions to estimate the foreground and background distributions. In this paper, we present a novel deep-learning-based algorithm which has much better understanding of objectness and can reduce user interactions to just a few clicks. Our algorithm transforms user-provided positive and negative clicks into two Euclidean distance maps which are then concatenated with the RGB channels of images to compose (image, user interactions) pairs. We generate many of such pairs by combining several random sampling strategies to model users' click patterns and use them to finetune deep Fully Convolutional Networks (FCNs). Finally the output probability maps of our FCN-8s model is integrated with graph cut optimization to refine the boundary segments. Our model is trained on the PASCAL segmentation dataset and evaluated on other datasets with different object classes. Experimental results on both seen and unseen objects demonstrate that our algorithm has a good generalization ability and is superior to all existing interactive object selection approaches.
Ning Xu 0007, Brian L. Price, Scott Cohen, Jimei Yang, Thomas S. Huang
CVPR2
2016 Object Contour Detection with a Fully Convolutional Encoder-Decoder Network
abstract
We develop a deep learning algorithm for contour detection with a fully convolutional encoder-decoder network. Different from previous low-level edge detection, our algorithm focuses on detecting higher-level object contours. Our network is trained end-to-end on PASCAL VOC with refined ground truth from inaccurate polygon annotations, yielding much higher precision in object contour detection than previous methods. We find that the learned model generalizes well to unseen object classes from the same supercategories on MS COCO and can match state-of-the-art edge detection on BSDS500 with fine-tuning. By combining with the multiscale combinatorial grouping algorithm, our method can generate high-quality segmented object proposals, which significantly advance the state-of-the-art on PASCAL VOC (improving average recall from 0.62 to 0.67) with a relatively small amount of candidates (~1660 per image).
Jimei Yang, Brian L. Price, Scott Cohen, Honglak Lee, Ming-Hsuan Yang 0001
CVPR2
2016 Unconstrained Salient Object Detection via Proposal Subset Optimization
abstract
We aim at detecting salient objects in unconstrained images. In unconstrained images, the number of salient objects (if any) varies from image to image, and is not given. We present a salient object detection system that directly outputs a compact set of detection windows, if any, for an input image. Our system leverages a Convolutional-Neural-Network model to generate location proposals of salient objects. Location proposals tend to be highly overlapping and noisy. Based on the Maximum a Posteriori principle, we propose a novel subset optimization framework to generate a compact set of detection windows out of noisy proposals. In experiments, we show that our subset optimization formulation greatly enhances the performance of our system, and our system attains 16-34% relative improvement in Average Precision compared with the state-of-the-art on three challenging salient object datasets.
Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
CVPR5
2016 SURGE: Surface Regularized Geometry Estimation from a Single Image
abstract
This paper introduces an approach to regularize 2.5D surface normal and depth predictions at each pixel given a single input image. The approach infers and reasons about the underlying 3D planar surfaces depicted in the image to snap predicted normals and depths to inferred planar surfaces, all while maintaining fine detail within objects. Our approach comprises two components: (i) a fourstream convolutional neural network (CNN) where depths, surface normals, and likelihoods of planar region and planar boundary are predicted at each pixel, followed by (ii) a dense conditional random field (DCRF) that integrates the four predictions such that the normals and depths are compatible with each other and regularized by the planar region and planar boundary information. The DCRF is formulated such that gradients can be passed to the surface normal and depth CNNs via backpropagation. In addition, we propose new planar wise metrics to evaluate geometry consistency within planar surfaces, which are more tightly related to dependent 3D editing applications. We show that our regularization yields a 30% relative improvement in planar consistency on the NYU v2 dataset.
Peng Wang 0001, Xiaohui Shen, Bryan C. Russell, Scott Cohen, Brian L. Price, Alan L. Yuille
NIPS5
2016 Automatic Portrait Segmentation for Image Stylization
abstract
Abstract Portraiture is a major art form in both photography and painting. In most instances, artists seek to make the subject stand out from its surrounding, for instance, by making it brighter or sharper. In the digital world, similar effects can be achieved by processing a portrait image with photographic or painterly filters that adapt to the semantics of the image. While many successful user‐guided methods exist to delineate the subject, fully automatic techniques are lacking and yield unsatisfactory results. Our paper first addresses this problem by introducing a new automatic segmentation algorithm dedicated to portraits. We then build upon this result and describe several portrait filters that exploit our automatic segmentation algorithm to generate high‐quality portraits.
Xiaoyong Shen, Aaron Hertzmann, Jiaya Jia, Sylvain Paris, Brian L. Price, Eli Shechtman, Ian Sachs
Comput. Graph. Forum5
2016 Discovering Primary Objects in Videos by Saliency Fusion and Iterative Appearance Estimation
abstract
In this paper, we propose a new method for detecting primary objects in unconstrained videos in a completely automatic setting. Here, we define the primary object in a video as the object that presents saliently in most of the frames. Unlike previous works considering only local saliency detection or common pattern discovery, the proposed method integrates the local visual/motion saliency extracted from each frame, global appearance consistency throughout the video, and spatiotemporal smoothness constraint on object trajectories. We first identify a temporal coherent salient region throughout the whole video, and then explicitly learn a global appearance model to distinguish the primary object against the background. In order to obtain high-quality saliency estimations from both appearance and motion cues, we propose a novel self-adaptive saliency map fusion method by learning the reliability of saliency maps from labeled data. As a whole, our method can robustly localize and track primary objects in diverse video content, and handle the challenges such as fast object and camera motion, large scale and appearance variation, background clutter, and pose deformation. Moreover, compared with some existing approaches that assume the object is present in all the frames, our approach can naturally handle the case where the object is present only in part of the frames, e.g., the object enters the scene in the middle of the video or leaves the scene before the video ends. We also propose a new video data set containing 51 videos for primary object detection with per-frame ground-truth labeling. Quantitative experiments on several challenging video data sets demonstrate the superiority of our method compared with the recent state of the arts.
Gangqiang Zhao, Junsong Yuan 0001, Xiaohui Shen, Zhe Lin 0001, Brian L. Price, Jonathan Brandt
IEEE Trans. Circuits Syst. Video Technol.6
2016 Fast Appearance Modeling for Automatic Primary Video Object Segmentation
abstract
Automatic segmentation of the primary object in a video clip is a challenging problem as there is no prior knowledge of the primary object. Most existing techniques thus adapt an iterative approach for foreground and background appearance modeling, i.e., fix the appearance model while optimizing the segmentation and fix the segmentation while optimizing the appearance model. However, these approaches may rely on good initialization and can be easily trapped in local optimal. In addition, they are usually time consuming for analyzing videos. To address these limitations, we propose a novel and efficient appearance modeling technique for automatic primary video object segmentation in the Markov random field (MRF) framework. It embeds the appearance constraint as auxiliary nodes and edges in the MRF structure, and can optimize both the segmentation and appearance model parameters simultaneously in one graph cut. The extensive experimental evaluations validate the superiority of the proposed approach over the state-of-the-art methods, in both efficiency and effectiveness.
Brian L. Price, Xiaohui Shen, Zhe Lin 0001, Junsong Yuan 0001
IEEE Trans. Image Process.2
2015 Effective learning-based illuminant estimation using simple features
abstract
Illumination estimation is the process of determining the chromaticity of the illumination in an imaged scene in order to remove undesirable color casts through white-balancing. While computational color constancy is a well-studied topic in computer vision, it remains challenging due to the ill-posed nature of the problem. One class of techniques relies on low-level statistical information in the image color distribution and works under various assumptions (e.g. Grey-World, White-Patch, etc). These methods have an advantage that they are simple and fast, but often do not perform well. More recent state-of-the-art methods employ learning-based techniques that produce better results, but often rely on complex features and have long evaluation and training times. In this paper, we present a learning-based method based on four simple color features and show how to use this with an ensemble of regression trees to estimate the illumination. We demonstrate that our approach is not only faster than existing learning-based methods in terms of both evaluation and training time, but also gives the best results reported to date on modern color constancy data sets.
Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown
CVPR2
2015 Towards unified depth and semantic prediction from a single image
abstract
Depth estimation and semantic segmentation are two fundamental problems in image understanding. While the two tasks are strongly correlated and mutually beneficial, they are usually solved separately or sequentially. Motivated by the complementary properties of the two tasks, we propose a unified framework for joint depth and semantic prediction. Given an image, we first use a trained Convolutional Neural Network (CNN) to jointly predict a global layout composed of pixel-wise depth values and semantic labels. By allowing for interactions between the depth and semantic information, the joint network provides more accurate depth prediction than a state-of-the-art CNN trained solely for depth prediction [6]. To further obtain fine-level details, the image is decomposed into local segments for region-level depth and semantic prediction under the guidance of global layout. Utilizing the pixel-wise global prediction and region-wise local prediction, we formulate the inference problem in a two-layer Hierarchical Conditional Random Field (HCRF) to produce the final depth and semantic map. As demonstrated in the experiments, our approach effectively leverages the advantages of both tasks and provides the state-of-the-art results.
Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille
CVPR5
2015 PatchCut: Data-driven object segmentation via local shape transfer
abstract
Object segmentation is highly desirable for image understanding and editing. Current interactive tools require a great deal of user effort while automatic methods are usually limited to images of special object categories or with high color contrast. In this paper, we propose a data-driven algorithm that uses examples to break through these limits. As similar objects tend to share similar local shapes, we match query image patches with example images in multiscale to enable local shape transfer. The transferred local shape masks constitute a patch-level segmentation solution space and we thus develop a novel cascade algorithm, PatchCut, for coarse-to-fine object segmentation. In each stage of the cascade, local shape mask candidates are selected to refine the estimated segmentation of the previous stage iteratively with color models. Experimental results on various datasets (Weizmann Horse, Fashionista, Object Discovery and PASCAL) demonstrate the effectiveness and robustness of our algorithm.
Jimei Yang, Brian L. Price, Scott Cohen, Zhe Lin 0001, Ming-Hsuan Yang 0001
CVPR2
2015 Salient Object Subitizing
abstract
People can immediately and precisely identify that an image contains 1, 2, 3 or 4 items by a simple glance. The phenomenon, known as Subitizing, inspires us to pursue the task of Salient Object Subitizing (SOS), i.e. predicting the existence and the number of salient objects in a scene using holistic cues. To study this problem, we propose a new image dataset annotated using an online crowdsourcing marketplace. We show that a proposed subitizing technique using an end-to-end Convolutional Neural Network (CNN) model achieves significantly better than chance performance in matching human labels on our dataset. It attains 94% accuracy in detecting the existence of salient objects, and 42–82% accuracy (chance is 20%) in predicting the number of salient objects (1, 2, 3, and 4+), without resorting to any object localization process. Finally, we demonstrate the usefulness of the proposed subitizing technique in two computer vision applications: salient object detection and object proposal.
Jianming Zhang 0001, Shugao Ma, Mehrnoosh Sameki, Stan Sclaroff, Margrit Betke, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
CVPR8
2015 Beyond White: Ground Truth Colors for Color Constancy Correction
abstract
A limitation in color constancy research is the inability to establish ground truth colors for evaluating corrected images. Many existing datasets contain images of scenes with a color chart included, however, only the chart's neutral colors (grayscale patches) are used to provide the ground truth for illumination estimation and correction. This is because the corrected neutral colors are known to lie along the achromatic line in the camera's color space (i.e. R=G=B), the correct RGB values of the other color patches are not known. As a result, most methods estimate a 3*3 diagonal matrix that ensures only the neutral colors are correct. In this paper, we describe how to overcome this limitation. Specifically, we show that under certain illuminations, a diagonal 3*3 matrix is capable of correcting not only neutral colors, but all the colors in a scene. This finding allows us to find the ground truth RGB values for the color chart in the camera's color space. We show how to use this information to correct all the images in existing datasets to have correct colors. Working from these new color corrected datasets, we describe how to modify existing color constancy algorithms to perform better image correction.
Dongliang Cheng, Brian L. Price, Scott Cohen, Michael S. Brown
ICCV2
2015 Joint Object and Part Segmentation Using Deep Learned Potentials
abstract
Segmenting semantic objects from images and parsing them into their respective semantic parts are fundamental steps towards detailed object understanding in computer vision. In this paper, we propose a joint solution that tackles semantic object and part segmentation simultaneously, in which higher object-level context is provided to guide part segmentation, and more detailed part-level localization is utilized to refine object segmentation. Specifically, we first introduce the concept of semantic compositional parts (SCP) in which similar semantic parts are grouped and shared among different objects. A two-stream fully convolutional network (FCN) is then trained to provide the SCP and object potentials at each pixel. At the same time, a compact set of segments can also be obtained from the SCP predictions of the network. Given the potentials and the generated segments, in order to explore long-range context, we finally construct an efficient fully connected conditional random field (FCRF) to jointly predict the final object and part labels. Extensive evaluation on three different datasets shows that our approach can mutually enhance the performance of object and part segmentation, and outperforms the current state-of-the-art on both tasks.
Peng Wang 0001, Xiaohui Shen, Zhe Lin 0001, Scott Cohen, Brian L. Price, Alan L. Yuille
ICCV5
2015 Minimum Barrier Salient Object Detection at 80 FPS
abstract
We propose a highly efficient, yet powerful, salient object detection method based on the Minimum Barrier Distance (MBD) Transform. The MBD transform is robust to pixel-value fluctuation, and thus can be effectively applied on raw pixels without region abstraction. We present an approximate MBD transform algorithm with 100X speedup over the exact algorithm. An error bound analysis is also provided. Powered by this fast MBD transform algorithm, the proposed salient object detection method runs at 80 FPS, and significantly outperforms previous methods with similar speed on four large benchmark datasets, and achieves comparable or better performance than state-of-the-art methods. Furthermore, a technique based on color whitening is proposed to extend our method to leverage the appearance-based backgroundness cue. This extended version further improves the performance, while still being one order of magnitude faster than all the other leading methods.
Jianming Zhang 0001, Stan Sclaroff, Zhe Lin 0001, Xiaohui Shen, Brian L. Price, Radomír Mech
ICCV5
2015 Inner and Inter Label Propagation: Salient Object Detection in the Wild
abstract
In this paper, we propose a novel label propagation-based method for saliency detection. A key observation is that saliency in an image can be estimated by propagating the labels extracted from the most certain background and object regions. For most natural images, some boundary superpixels serve as the background labels and the saliency of other superpixels are determined by ranking their similarities to the boundary labels based on an inner propagation scheme. For images of complex scenes, we further deploy a threecue-center-biased objectness measure to pick out and propagate foreground labels. A co-transduction algorithm is devised to fuse both boundary and objectness labels based on an inter propagation scheme. The compactness criterion decides whether the incorporation of objectness labels is necessary, thus greatly enhancing computational efficiency. Results on five benchmark data sets with pixelwise accurate annotations show that the proposed method achieves superior performance compared with the newest state-of-the-arts in terms of different evaluation metrics.
Hongyang Li 0001, Huchuan Lu, Zhe Lin 0001, Xiaohui Shen, Brian L. Price
IEEE Trans. Image Process.5
2015 Adaptive Metric Learning for Saliency Detection
abstract
In this paper, we propose a novel adaptive metric learning algorithm (AML) for visual saliency detection. A key observation is that the saliency of a superpixel can be estimated by the distance from the most certain foreground and background seeds. Instead of measuring distance on the Euclidean space, we present a learning method based on two complementary Mahalanobis distance metrics: 1) generic metric learning (GML) and 2) specific metric learning (SML). GML aims at the global distribution of the whole training set, while SML considers the specific structure of a single image. Considering that multiple similarity measures from different views may enhance the relevant information and alleviate the irrelevant one, we try to fuse the GML and SML together and experimentally find the combining result does work well. Different from the most existing methods which are directly based on low-level features, we devise a superpixelwise Fisher vector coding approach to better distinguish salient objects from the background. We also propose an accurate seeds selection mechanism and exploit contextual and multiscale information when constructing the final saliency map. Experimental results on various image sets show that the proposed AML performs favorably against the state-of-the-arts.
Huchuan Lu, Zhe Lin 0001, Xiaohui Shen, Brian L. Price
IEEE Trans. Image Process.5
2014 Semantic Object Selection
abstract
Interactive object segmentation has great practical importance in computer vision. Many interactive methods have been proposed utilizing user input in the form of mouse clicks and mouse strokes, and often requiring a lot of user intervention. In this paper, we present a system with a far simpler input method: the user needs only give the name of the desired object. With the tag provided by the user we do a text query of an image database to gather exemplars of the object. Using object proposals and borrowing ideas from image retrieval and object detection, the object is localized in the target image. An appearance model generated from the exemplars and the location prior are used in an energy minimization framework to select the object. Our method outperforms the state-of-the-art on existing datasets and on a more challenging dataset we collected.
Ejaz Ahmed 0002, Scott Cohen, Brian L. Price
CVPR3
2014 Context Driven Scene Parsing with Attention to Rare Classes
abstract
This paper presents a scalable scene parsing algorithm based on image retrieval and superpixel matching. We focus on rare object classes, which play an important role in achieving richer semantic understanding of visual scenes, compared to common background classes. Towards this end, we make two novel contributions: rare class expansion and semantic context description. First, considering the long-tailed nature of the label distribution, we expand the retrieval set by rare class exemplars and thus achieve more balanced superpixel classification results. Second, we incorporate both global and local semantic context information through a feedback based mechanism to refine image retrieval and superpixel matching. Results on the SIFTflow and LMSun datasets show the superior performance of our algorithm, especially on the rare classes, without sacrificing overall labeling accuracy.
Jimei Yang, Brian L. Price, Scott Cohen, Ming-Hsuan Yang 0001
CVPR2
2014 Depth-based patch scaling for content-aware stereo image completion
abstract
A number of recent algorithms have been proposed for working with stereo image pairs in ways that are already familiar to users of single-image editing tools. In particular, Morse, et al. (2012) have proposed a method for performing image completion in stereo images so as to maintain stereoscopic consistency. Like prior work in stereo completion, this method drew source texture only from regions at the same depth as the target region, which while helping the result can sometimes overly limit the pool of suitable source textures. Other methods such as the Generalized PatchMatch approach of Barnes, et al. (2010) have used scaled (and otherwise transformed) source texture to improve the quality of the completed target region, but these methods rely on randomly sampling the scale (or transformation) space without knowledge of scene geometry. This paper extends stereo image completion to include source textures scaled according to the relative differences in depth between image regions. Limited random sampling is used to make the method robust to minor errors in the stereo disparities and to provide for non-uniform aspect ratios, but with far fewer random samples than prior unrestrained sampling of scale. A preference for unscaled or downsampled source textures rather than upsampled ones is incorporated into the objective function and avoids an inherent matching bias towards low-frequency regions. Results demonstrate that using scene geometry to drive scale selection results in improved image completion compared to either single-image completion or prior methods for stereo completion.
Joel Howard, Bryan S. Morse, Scott Cohen, Brian L. Price
WACV4
2014 Temporally coherent and spatially accurate video matting
abstract
Abstract Image and video matting are still challenging problems in areas with low foreground‐background contrast. Video matting also has the challenge of ensuring temporally coherent mattes because the human visual system is highly sensitive to temporal jitter and flickering. On the other hand, video provides the opportunity to use information from other frames to improve the matte accuracy on a given frame. In this paper, we present a new video matting approach that improves the temporal coherence while maintaining high spatial accuracy in the computed mattes. We build sample sets of temporal and local samples that cover all the color distributions of the object and background over all previous frames. This helps guarantee spatial accuracy and temporal coherence by ensuring that proper samples are found even when distantly located in space or time. An explicit energy term encourages temporal consistency in the mattes derived from the selected samples. In addition, we use localized texture features to improve spatial accuracy in low contrast regions where color distributions overlap. The proposed method results in better spatial accuracy and temporal coherence than existing video matting methods.
E. Shahrian, Brian L. Price, Scott Cohen, D. Rajan
Comput. Graph. Forum2
2013 Stereo+Kinect for High Resolution Stereo Correspondences
abstract
In this work, we combine the complementary depth sensors Kinect and stereo image matching to obtain high quality correspondences. Our goal is to obtain a dense disparity map at the spatial and depth resolution of the stereo cameras (4-12 MP). We propose a global optimization scheme, where both the data and smoothness costs are derived using sensor confidences and low resolution geometry from Kinect. A spatially varying search range is used to limit the number of potential disparities at each pixel. The smoothness prior is Based on available low resolution depth from Kinect rather than image gradients, thus performing better in both textured areas with smooth depth and texture-less areas with depth gradient. We also propose a spatially varying smoothness weight to better handle occlusion areas, and the relative contribution of the two energy terms. We demonstrate how the two sensors can be effectively fused to obtain correct scene depth in ambiguous areas, as well as fine structural details in textured areas.
Gowri Somanath, Scott Cohen, Brian L. Price, Chandra Kambhamettu
3DV3
2013 High-Quality Stereo Video Matching via User Interaction and Space-Time Propagation
abstract
Even current state-of-the-art automatic stereo matching methods often struggle on natural images and videos, in great part due to fundamental matching ambiguities in low texture regions and a lack of higher level object knowledge. Stereo image matching can benefit greatly from user input to guide the matching process and help disambiguate matches. Applying interactive correction tools from scratch on each frame of a video would not only be throwing away valuable information provided by the user on other frames, but would also likely be too time consuming to be practical for video even if excellent disparity results could be obtained within a few minutes on each frame. In this work, we propose a stereo video matching system that allows user interaction to obtain high quality, dense disparity maps on key frames and then intelligently propagates the user input and key frame disparities to automatically produce high quality disparity maps on intermediate frames. The disparity maps on key frames are obtained using several novel, easy-to-use, and effective interactive tools. Our novel propagation algorithm estimates 3D transformations that map user corrected areas in key frames to intermediate frames. Experiments demonstrate the effectiveness and efficiency of our hybrid interactive/automatic approach.
Brian L. Price, Scott Cohen, Ruigang Yang
3DV2
2013 Improving Image Matting Using Comprehensive Sampling Sets
abstract
In this paper, we present a new image matting algorithm that achieves state-of-the-art performance on a benchmark dataset of images. This is achieved by solving two major problems encountered by current sampling based algorithms. The first is that the range in which the foreground and background are sampled is often limited to such an extent that the true foreground and background colors are not present. Here, we describe a method by which a more comprehensive and representative set of samples is collected so as not to miss out on the true samples. This is accomplished by expanding the sampling range for pixels farther from the foreground or background boundary and ensuring that samples from each color distribution are included. The second problem is the overlap in color distributions of foreground and background regions. This causes sampling based methods to fail to pick the correct samples for foreground and background. Our design of an objective function forces those foreground and background samples to be picked that are generated from well-separated distributions. Comparison on the dataset at and evaluation by www.alphamatting.com shows that the proposed method ranks first in terms of error measures used in the website.
Ehsan Shahrian, Deepu Rajan, Brian L. Price, Scott Cohen
CVPR3
2011 StereoCut: Consistent interactive object selection in stereo image pairs
abstract
Methods of interacting with stereo image pairs are important for handling the increasing amount of stereoscopic 3D data now being produced. In this paper, we introduce a framework for interactively selecting objects in two stereo images simultaneously using graph cut. A key contribution of our method is the use of stereo correspondence probability distributions to govern the strength of the connection between the two images. This allows information from arbitrary stereo matching algorithms to be utilized by our method. We show how to enforce consistency in these distributions to improve the results. For comparisons, we introduce a new dataset of stereo images and ground truth selections. We evaluate different correspondence distributions and show that our method is effective in selecting objects from stereo pairs.
Brian L. Price, Scott Cohen
ICCV1
2010 Simultaneous foreground, background, and alpha estimation for image matting
abstract
Image matting is the process of extracting a soft segmentation of an object in an image as defined by the matting equation. Most current techniques focus largely on computing the alpha values of unknown pixels and treat computation of the foreground and background colors as an afterthought, if at all. However, for many applications, such as compositing an object into a new scene or deleting an object from the scene, the foreground and background colors are vital for an acceptable answer. We propose a method of solving for the foreground, background, and alpha of an unknown region in an image simultaneously. This allows for novel constraints to be placed directly on the foreground and background as well as on alpha. We show through both visual results and quantitative measurements on standard datasets that this approach produces more accurate foreground and background values at each pixel while maintaining competitive results on the alpha matte.
Brian L. Price, Bryan S. Morse, Scott Cohen
CVPR1
2010 Geodesic graph cut for interactive image segmentation
abstract
Interactive segmentation is useful for selecting objects of interest in images and continues to be a topic of much study. Methods that grow regions from foreground/background seeds, such as the recent geodesic segmentation approach, avoid the boundary-length bias of graph-cut methods but have their own bias towards minimizing paths to the seeds, resulting in increased sensitivity to seed placement. The lack of edge modeling in geodesic or similar approaches limits their ability to precisely localize object boundaries, something at which graph-cut methods generally excel. This paper presents a method for combining geodesic-distance information with edge information in a graphcut optimization framework, leveraging the complementary strengths of each. Rather than a fixed combination we use the distinctiveness of the foreground/background color models to predict the effectiveness of the geodesic distance term and adjust the weighting accordingly. We also introduce a spatially varying weighting that decreases the potential for shortcutting in object interiors while transferring greater control to the edge term for better localization near object boundaries. Results show our method is less prone to shortcutting than typical graph cut methods while being less sensitive to seed placement and better at edge localization than geodesic methods. This leads to increased segmentation accuracy and reduced effort on the part of the user.
Brian L. Price, Bryan S. Morse, Scott Cohen
CVPR1
2010 Color Adjacency Modeling for Improved Image and Video Segmentation
abstract
Color models are often used for representing object appearance for foreground segmentation applications. The relationships between colors can be just as useful for object selection. In this paper, we present a method of modeling color adjacency relationships. By using color adjacency models, the importance of an edge in a given application can be determined and scaled accordingly. We apply our model to foreground segmentation of similar images and video. We show that given one previously-segmented image, we can greatly reduce the error when automatically segmenting other images by using our color adjacency model to weight the likelihood that an edge is part of the desired object boundary.
Brian L. Price, Bryan S. Morse, Scott Cohen
ICPR1
2009 LIVEcut: Learning-based interactive video segmentation by evaluation of multiple propagated cues
abstract
Video sequences contain many cues that may be used to segment objects in them, such as color, gradient, color adjacency, shape, temporal coherence, camera and object motion, and easily-trackable points. This paper introduces LIVEcut, a novel method for interactively selecting objects in video sequences by extracting and leveraging as much of this information as possible. Using a graph-cut optimization framework, LIVEcut propagates the selection forward frame by frame, allowing the user to correct any mistakes along the way if needed. Enhanced methods of extracting many of the features are provided. In order to use the most accurate information from the various potentially-conflicting features, each feature is automatically weighted locally based on its estimated accuracy using the previous implicitly-validated frame. Feature weights are further updated by learning from the user corrections required in the previous frame. The effectiveness of LIVEcut is shown through timing comparisons to other interactive methods, accuracy comparisons to unsupervised methods, and qualitatively through selections on various video sequences.
Brian L. Price, Bryan S. Morse, Scott Cohen
ICCV1
2007 Interactive segmentation of image volumes with Live Surface
Christopher J. Armstrong, Brian L. Price, Bill Barrett
Comput. Graph.2
2006 Object-based vectorization for interactive image editing
Brian L. Price, Bill Barrett
Vis. Comput.1