Aseem Agarwala

dblp:91/28 · DBLP profile ↗
← Back
43ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 5 first-author · 1 since 2021Artificial intelligence and machine learning · 11 · 1 since 2021Human-computer interaction and ubiquitous computing · 4

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
28 papers
Image and video processing · 29% Visual content generation and editing · 17% Rendering · 14%
Artificial intelligence
11 papers
3D vision · 27% Learning paradigms · 14% Efficient and distributed learning · 13%
Human-computer interaction and pervasive computing
6 papers
User interface design and tools · 43% Haptics and multimodal interaction · 29% Human-AI interaction · 14%

Topics — the 30 heaviest of 85, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Geometric modeling and processing
3d segmentation
0.612022
Neural Volumetric Object Selection · CVPR 2022
Rendering
neural radiance fields
0.612022
Neural Volumetric Object Selection · CVPR 2022
Computer vision › 3D vision › multi-view geometry
homography estimation
0.412020
Deep Homography Estimation for Dynamic Scenes · CVPR 2020
Machine learning › Learning paradigms
multi-task learning
0.412020
Deep Homography Estimation for Dynamic Scenes · CVPR 2020
Computer vision › Face, body and person analysis
facial expression analysis
0.412019
A Compact Embedding for Facial Expression Similarity · CVPR 2019
Machine learning › Representation and self-supervised learning › representation learning › metric learning
similarity embedding
0.412019
A Compact Embedding for Facial Expression Similarity · CVPR 2019
Image and video processing
video stabilization
0.332011
Subspace video stabilization · ACM Trans. Graph. 2011
Content-preserving warps for 3D video stabilization · ACM Trans. Graph. 2009
Light field video stabilization · ICCV 2009
Visual content generation and editing › image editing
image compositing
0.352012
Understanding and improving the realism of image composites · ACM Trans. Graph. 2012
Efficient gradient-domain compositing using quadtrees · ACM Trans. Graph. 2007
Interactive digital photomontage · ACM Trans. Graph. 2004
Image and video processing › video frame interpolation
video frame synthesis
0.312017
Video Frame Synthesis Using Deep Voxel Flow · ICCV 2017
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.212015
DeepFont: Identify Your Font from An Image · ACM Multimedia 2015
Machine learning › Efficient and distributed learning
model compression
0.212015
DeepFont: Identify Your Font from An Image · ACM Multimedia 2015
Image and video processing › document image analysis
font recognition
0.212015
DeepFont: A System for Font Recognition and Similarity · ACM Multimedia 2015
Natural language and speech › Information extraction and text analysis › document understanding › document image analysis
font recognition
0.212014
Large-Scale Visual Font Recognition · CVPR 2014
Visual content generation and editing
graphic design
0.212014
Learning Layouts for Single-PageGraphic Designs · IEEE Trans. Vis. Comput. Graph. 2014
Visual content generation and editing
layout generation
0.212014
Learning Layouts for Single-PageGraphic Designs · IEEE Trans. Vis. Comput. Graph. 2014
Computational photography and imaging
image-based modeling
0.212013
Image-Based Remodeling · IEEE Trans. Vis. Comput. Graph. 2013
Rendering
photorealistic rendering
0.212013
Image-Based Remodeling · IEEE Trans. Vis. Comput. Graph. 2013
Rendering › texture mapping
view-dependent texture mapping
0.212013
Image-Based Remodeling · IEEE Trans. Vis. Comput. Graph. 2013
Haptics and multimodal interaction
multimodal interaction
0.212013
PixelTone: a multimodal interface for image editing · CHI 2013
Human-AI interaction › large language model interaction › language-based interaction
natural language interface
0.212013
PixelTone: a multimodal interface for image editing · CHI 2013
Haptics and multimodal interaction › multimodal interaction
speech and sketch input
0.212013
PixelTone: a multimodal interface for image editing · CHI 2013
Rendering
realism assessment
0.112012
Understanding and improving the realism of image composites · ACM Trans. Graph. 2012
Visual content generation and editing › video editing
video composition
0.112012
Selectively de-animating video · ACM Trans. Graph. 2012
Visual content generation and editing
video editing
0.112012
Selectively de-animating video · ACM Trans. Graph. 2012
Image and video processing › color image processing › color image analysis
color compatibility
0.112011
Color compatibility from large datasets · ACM Trans. Graph. 2011
Visualization and visual analytics › visual encoding
color design
0.112011
Color compatibility from large datasets · ACM Trans. Graph. 2011
Multimedia analysis and retrieval › multimedia analysis › multimedia collection analysis › image collection analysis
photo selection
0.112011
Candid portrait selection from video · ACM Trans. Graph. 2011
Image and video processing › video stabilization
trajectory smoothing
0.112011
Subspace video stabilization · ACM Trans. Graph. 2011
Image and video processing
image warping
0.112010
Image warps for artistic perspective manipulation · ACM Trans. Graph. 2010
Image and video processing › image warping
perspective manipulation
0.112010
Image warps for artistic perspective manipulation · ACM Trans. Graph. 2010

Methods — techniques the papers use, named apart from their topics

convolutional neural network · 0.7voxel feature embedding · 0.6user scribbles · 0.6multi-view image features · 0.6deep neural network · 0.6stacked convolutional auto-encoder · 0.4model compression · 0.4multi-task learning · 0.4multi-scale neural network · 0.4user study · 0.4triplet ranking · 0.4neural network embedding · 0.4graph-cut optimization · 0.2structure from motion · 0.2local feature metric learning · 0.2local feature embedding · 0.2font similarity metric learning · 0.2crowdsourcing · 0.2
YearPublicationVenuePosition
2022 Neural Volumetric Object Selection
abstract
We introduce an approach for selecting objects in neural volumetric 3D representations, such as multi-plane images (MPI) and neural radiance fields (NeRF). Our approach takes a set of foreground and background 2D user scribbles in one view and automatically estimates a 3D segmentation of the desired object, which can be rendered into novel views. To achieve this result, we propose a novel voxel feature embedding that incorporates the neural volumetric 3D representation and multi-view image features from all input views. To evaluate our approach, we introduce a new dataset of human-provided segmentation masks for depicted objects in real-world multi-view scene captures. We show that our approach out-performs strong baselines, including 2D segmentation and 3D segmentation approaches adapted to our task.
Zhongzheng Ren, Aseem Agarwala, Bryan C. Russell, Alexander G. Schwing, Oliver Wang
CVPR2
2020 Deep Homography Estimation for Dynamic Scenes
abstract
Homography estimation is an important step in many computer vision problems. Recently, deep neural network methods have shown to be favorable for this problem when compared to traditional methods. However, these new methods do not consider dynamic content in input images. They train neural networks with only image pairs that can be perfectly aligned using homographies. This paper investigates and discusses how to design and train a deep neural network that handles dynamic scenes. We first collect a large video dataset with dynamic content. We then develop a multi-scale neural network and show that when properly trained using our new dataset, this neural network can already handle dynamic scenes to some extent. To estimate a homography of a dynamic scene in a more principled way, we need to identify the dynamic content. Since dynamic content detection and homography estimation are two tightly coupled tasks, we follow the multi-task learning principles and augment our multi-scale network such that it jointly estimates the dynamics masks and homographies. Our experiments show that our method can robustly estimate homography for challenging scenarios with dynamic scenes, blur artifacts, or lack of textures.
Hoang Le, Feng Liu 0015, Aseem Agarwala
CVPR4
2019 A Compact Embedding for Facial Expression Similarity
abstract
Most of the existing work on automatic facial expression analysis focuses on discrete emotion recognition, or facial action unit detection. However, facial expressions do not always fall neatly into pre-defined semantic categories. Also, the similarity between expressions measured in the action unit space need not correspond to how humans perceive expression similarity. Different from previous work, our goal is to describe facial expressions in a continuous fashion using a compact embedding space that mimics human visual preferences. To achieve this goal, we collect a large-scale faces-in-the-wild dataset with human annotations in the form: Expressions A and B are visually more similar when compared to expression C, and use this dataset to train a neural network that produces a compact (16-dimensional) expression embedding. We experimentally demonstrate that the learned embedding can be successfully used for various applications such as expression retrieval, photo album summarization, and emotion recognition. We also show that the embedding learned using the proposed dataset performs better than several other embeddings learned using existing emotion or action unit datasets.
Raviteja Vemulapalli, Aseem Agarwala
CVPR2
2017 Video Frame Synthesis Using Deep Voxel Flow
abstract
We address the problem of synthesizing new video frames in an existing video, either in-between existing frames (interpolation), or subsequent to them (extrapolation). This problem is challenging because video appearance and motion can be highly complex. Traditional optical-flow-based solutions often fail where flow estimation is challenging, while newer neural-network-based methods that hallucinate pixel values directly often produce blurry results. We combine the advantages of these two methods by training a deep network that learns to synthesize video frames by flowing pixel values from existing ones, which we call deep voxel flow. Our method requires no human supervision, and any video can be used as training data by dropping, and then learning to predict, existing frames. The technique is efficient, and can be applied at any video resolution. We demonstrate that our method produces results that both quantitatively and qualitatively improve upon the state-of-the-art.
Ziwei Liu 0002, Raymond A. Yeh, Xiaoou Tang, Yiming Liu 0001, Aseem Agarwala
ICCV5
2017 Style-based exploration of illustration datasets
Elena Garces 0001, Aseem Agarwala, Aaron Hertzmann, Diego Gutierrez
Multim. Tools Appl.2
2015 DesignScape: Design with Interactive Layout Suggestions
abstract
Creating graphic designs can be challenging for novice users. This paper presents DesignScape, a system which aids the design process by making interactive layout suggestions, i.e., changes in the position, scale, and alignment of elements. The system uses two distinct but complementary types of suggestions: refinement suggestions, which improve the current layout, and brainstorming suggestions, which change the style. We investigate two interfaces for interacting with suggestions. First, we develop a suggestive interface, where suggestions are previewed and can be accepted. Second, we develop an adaptive interface where elements move automatically to improve the layout. We compare both interfaces with a baseline without suggestions, and show that for novice designers, both interfaces produce significantly better layouts, as evaluated by other novices.
Peter O'Donovan, Aseem Agarwala, Aaron Hertzmann
CHI2
2015 DeepFont: A System for Font Recognition and Similarity
abstract
We develop the DeepFont system, a large-scale learning-based solution for automatic font identification, organization and selection. In this proposed technical demonstration, we will give our audience a tour to the DeepFont system, with the focus on its impacts on real consumer products, including but not limited to: 1) a cloud-based iOS App for font recognition; 2) a web-based tool for font similarity evaluation and discovery.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Yuyan Song, Joseph Hsieh, Sarah Kong, Thomas S. Huang
ACM Multimedia6
2015 DeepFont: Identify Your Font from An Image
abstract
As font is one of the core design concepts, automatic font identification and similar font suggestion from an image or photo has been on the wish list of many designers. We study the Visual Font Recognition (VFR) problem [4] LFE, and advance the state-of-the-art remarkably by developing the DeepFont system. First of all, we build up the first available large-scale VFR dataset, named AdobeVFR, consisting of both labeled synthetic data and partially labeled real-world data. Next, to combat the domain mismatch between available training and testing data, we introduce a Convolutional Neural Network (CNN) decomposition approach, using a domain adaptation technique based on a Stacked Convolutional Auto-Encoder (SCAE) that exploits a large corpus of unlabeled real-world text images combined with synthetic data preprocessed in a specific way. Moreover, we study a novel learning-based model compression approach, in order to reduce the DeepFont model size without sacrificing its performance. The DeepFont system achieves an accuracy of higher than 80% (top-5) on our collected dataset, and also produces a good font similarity measure for font selection and suggestion. We also achieve around 6 times compression of the model without any visible loss of recognition accuracy.
Zhangyang Wang, Jianchao Yang, Hailin Jin, Eli Shechtman, Aseem Agarwala, Jonathan Brandt, Thomas S. Huang
ACM Multimedia5
2014 Recognizing Image Style
Sergey Karayev, Matthew Trentacoste, Helen Han, Aseem Agarwala, Trevor Darrell, Aaron Hertzmann, Holger Winnemöller
BMVC4
2014 Large-Scale Visual Font Recognition
abstract
This paper addresses the large-scale visual font recognition (VFR) problem, which aims at automatic identification of the typeface, weight, and slope of the text in an image or photo without any knowledge of content. Although visual font recognition has many practical applications, it has largely been neglected by the vision community. To address the VFR problem, we construct a large-scale dataset containing 2,420 font classes, which easily exceeds the scale of most image categorization datasets in computer vision. As font recognition is inherently dynamic and open-ended, i.e., new classes and data for existing categories are constantly added to the database over time, we propose a scalable solution based on the nearest class mean classifier (NCM). The core algorithm is built on local feature embedding, local feature metric learning and max-margin template selection, which is naturally amenable to NCM and thus to such open-ended classification problems. The new algorithm can generalize to new classes and new data at little added cost. Extensive experiments demonstrate that our approach is very effective on our synthetic test images, and achieves promising results on real world test images.
Jianchao Yang, Hailin Jin, Jonathan Brandt, Eli Shechtman, Aseem Agarwala, Tony X. Han
CVPR6
2014 User-Assisted Video Stabilization
abstract
Abstract We present a user‐assisted video stabilization algorithm that is able to stabilize challenging videos when state‐of‐the‐art automatic algorithms fail to generate a satisfactory result. Current methods do not give the user any control over the look of the final result. Users either have to accept the stabilized result as is, or discard it should the stabilization fail to generate a smooth output. Our system introduces two new modes of interaction that allow the user to improve the unsatisfactory stabilized video. First, we cluster tracks and visualize them on the warped video. The user ensures that appropriate tracks are selected by clicking on track clusters to include or exclude them. Second, the user can directly specify how regions in the output video should look by drawing quadrilaterals to select and deform parts of the frame. These user‐provided deformations reduce undesirable distortions in the video. Our algorithm then computes a stabilized video using the user‐selected tracks, while respecting the user‐modified regions. The process of interactively removing user‐identified artifacts can sometimes introduce new ones, though in most cases there is a net improvement. We demonstrate the effectiveness of our system with a variety of challenging hand held videos.
Jiamin Bai, Aseem Agarwala, Maneesh Agrawala, Ravi Ramamoorthi
Comput. Graph. Forum2
2014 A similarity measure for illustration style
abstract
This paper presents a method for measuring the similarity in style between two pieces of vector art, independent of content. Similarity is measured by the differences between four types of features: color, shading, texture, and stroke. Feature weightings are learned from crowdsourced experiments. This perceptual similarity enables style-based search. Using our style-based search feature, we demonstrate an application that allows users to create stylistically-coherent clip art mash-ups.
Elena Garces 0001, Aseem Agarwala, Diego Gutierrez, Aaron Hertzmann
ACM Trans. Graph.2
2014 Exploratory font selection using crowdsourced attributes
abstract
This paper presents interfaces for exploring large collections of fonts for design tasks. Existing interfaces typically list fonts in a long, alphabetically-sorted menu that can be challenging and frustrating to explore. We instead propose three interfaces for font selection. First, we organize fonts using high-level descriptive attributes, such as "dramatic" or "legible." Second, we organize fonts in a tree-based hierarchical menu based on perceptual similarity. Third, we display fonts that are most similar to a user's currently-selected font. These tools are complementary; a user may search for "graceful" fonts, select a reasonable one, and then refine the results from a list of fonts similar to the selection. To enable these tools, we use crowdsourcing to gather font attribute data, and then train models to predict attribute values for new fonts. We use attributes to help learn a font similarity metric using crowdsourced comparisons. We evaluate the interfaces against a conventional list interface and find that our interfaces are preferred to the baseline. Our interfaces also produce better results in two real-world tasks: finding the nearest match to a target font, and font selection for graphic designs.
Peter O'Donovan, Janis Libeks, Aseem Agarwala, Aaron Hertzmann
ACM Trans. Graph.3
2014 Learning Layouts for Single-PageGraphic Designs
abstract
This paper presents an approach for automatically creating graphic design layouts using a new energy-based model derived from design principles. The model includes several new algorithms for analyzing graphic designs, including the prediction of perceived importance, alignment detection, and hierarchical segmentation. Given the model, we use optimization to synthesize new layouts for a variety of single-page graphic designs. Model parameters are learned with Nonlinear Inverse Optimization (NIO) from a small number of example layouts. To demonstrate our approach, we show results for applications including generating design layouts in various styles, retargeting designs to new sizes, and improving existing designs. We also compare our automatic results with designs created using crowdsourcing and show that our approach performs slightly better than novice designers.
Peter O'Donovan, Aseem Agarwala, Aaron Hertzmann
IEEE Trans. Vis. Comput. Graph.2
2013 PixelTone: a multimodal interface for image editing
abstract
Photo editing can be a challenging task, and it becomes even more difficult on the small, portable screens of mobile devices that are now frequently used to capture and edit images. To address this problem we present PixelTone, a multimodal photo editing interface that combines speech and direct manipulation. We observe existing image editing practices and derive a set of principles that guide our design. In particular, we use natural language for expressing desired changes to an image, and sketching to localize these changes to specific regions. To support the language commonly used in photo-editing we develop a customized natural language interpreter that maps user phrases to specific image processing operations. Finally, we perform a user study that evaluates and demonstrates the effectiveness of our interface.
Gierad Laput, Mira Dontcheva, Gregg Wilensky, Walter Chang, Aseem Agarwala, Jason Linder, Eytan Adar
CHI5
2013 Automatic Cinemagraph Portraits
abstract
Abstract Cinemagraphs are a popular new type of visual media that lie in‐between photos and video; some parts of the frame are animated and loop seamlessly, while other parts of the frame remain completely still. Cinemagraphs are especially effective for portraits because they capture the nuances of our dynamic facial expressions. We present a completely automatic algorithm for generating portrait cinemagraphs from a short video captured with a hand‐held camera. Our algorithm uses a combination of face tracking and point tracking to segment face motions into two classes: gross, large‐scale motions that should be removed from the video, and dynamic facial expressions that should be preserved. This segmentation informs a spatially‐varying warp that removes the large‐scale motion, and a graph‐cut segmentation of the frame into dynamic and still regions that preserves the finer‐scale facial expression motions. We demonstrate the success of our method with a variety of results and a comparison to previous work.
Jiamin Bai, Aseem Agarwala, Maneesh Agrawala, Ravi Ramamoorthi
Comput. Graph. Forum2
2013 Learning and Applying Color Styles From Feature Films
abstract
Abstract Directors employ a process called “color grading” to add color styles to feature films. Color grading is used for a number of reasons, such as accentuating a certain emotion or expressing the signature look of a director. We collect a database of feature film clips and label them with tags such as director, emotion, and genre. We then learn a model that maps from the low‐level color and tone properties of film clips to the associated labels. This model allows us to examine a number of common hypotheses on the use of color to achieve goals, such as specific emotions. We also describe a method to apply our learned color styles to new images and videos. Along with our analysis of color grading techniques, we demonstrate a number of images and videos that are automatically filtered to resemble certain film styles.
Su Xue, Aseem Agarwala, Julie Dorsey, Holly E. Rushmeier
Comput. Graph. Forum2
2013 Image-Based Remodeling
abstract
Imagining what a proposed home remodel might look like without actually performing it is challenging. We present an image-based remodeling methodology that allows real-time photorealistic visualization during both the modeling and remodeling process of a home interior. Large-scale edits, like removing a wall or enlarging a window, are performed easily and in real time, with realistic results. Our interface supports the creation of concise, parameterized, and constrained geometry, as well as remodeling directly from within the photographs. Real-time texturing of modified geometry is made possible by precomputing view-dependent textures for all faces that are potentially visible to each original camera viewpoint, blending multiple viewpoints and hole-filling when necessary. The resulting textures are stored and accessed efficiently enabling intuitive real-time realistic visualization, modeling, and editing of the building interior.
Alex Colburn, Aseem Agarwala, Aaron Hertzmann, Brian Curless, Michael F. Cohen
IEEE Trans. Vis. Comput. Graph.2
2012 Selectively de-animating video
abstract
We present a semi-automated technique for selectivelydeanimatingvideo to remove the large-scale motions of one or more objects so that other motions are easier to see. The user draws strokes to indicate the regions of the video that should be immobilized, and our algorithm warps the video to remove the large-scale motion of these regions while leaving finer-scale, relative motions intact. However, such warps may introduce unnatural motions in previously motionless areas, such as background regions. We therefore use a graph-cut-based optimization to composite the warped video regions with still frames from the input video; we also optionally loop the output in a seamless manner. Our technique enables a number of applications such as clearer motion visualization, simpler creation of artisticcinemagraphs(photos that include looping motions in some regions), and new ways to edit appearance and complicated motion paths in video by manipulating a de-animated representation. We demonstrate the success of our technique with a number of motion visualizations, cinemagraphs and video editing examples created from a variety of short input videos, as well as visual and numerical comparison to previous techniques.
Jiamin Bai, Aseem Agarwala, Maneesh Agrawala, Ravi Ramamoorthi
ACM Trans. Graph.2
2012 Understanding and improving the realism of image composites
abstract
Compositing is one of the most commonly performed operations in computer graphics. A realistic composite requires adjusting the appearance of the foreground and background so that they appear compatible; unfortunately, this task is challenging and poorly understood. We use statistical and visual perception experiments to study the realism of image composites. First, we evaluate a number of standard 2D image statistical measures, and identify those that are most significant in determining the realism of a composite. Then, we perform a human subjects experiment to determine how the changes in these key statistics influence human judgements of composite realism. Finally, we describe a data-driven algorithm that automatically adjusts these statistical measures in a foreground to make it more compatible with its background in a composite. We show a number of compositing results, and evaluate the performance of both our algorithm and previous work with a human subjects study.
Su Xue, Aseem Agarwala, Julie Dorsey, Holly E. Rushmeier
ACM Trans. Graph.2
2011 Candid portrait selection from video
abstract
In this paper, we train a computer to select still frames from video that work well as candid portraits. Because of the subjective nature of this task, we conduct a human subjects study to collect ratings of video frames across multiple videos. Then, we compute a number of features and train a model to predict the average rating of a video frame. We evaluate our model with cross-validation, and show that it is better able to select quality still frames than previous techniques, such as simply omitting frames that contain blinking or motion blur, or selecting only smiles. We also evaluate our technique qualitatively on videos that were not part of our validation set, and were taken outdoors and under different lighting conditions.
Juliet Fiss, Aseem Agarwala, Brian Curless
ACM Trans. Graph.2
2011 Subspace video stabilization
abstract
We present a robust and efficient approach to video stabilization that achieves high-quality camera motion for a wide range of videos. In this article, we focus on the problem of transforming a set of input 2D motion trajectories so that they are both smooth and resemble visually plausible views of the imaged scene; our key insight is that we can achieve this goal by enforcing subspace constraints on feature trajectories while smoothing them. Our approach assembles tracked features in the video into a trajectory matrix, factors it into two low-rank matrices, and performs filtering or curve fitting in a low-dimensional linear space. In order to process long videos, we propose a moving factorization that is both efficient and streamable. Our experiments confirm that our approach can efficiently provide stabilization results comparable with prior 3D methods in cases where those methods succeed, but also provides smooth camera motions in cases where such approaches often fail, such as videos that lack parallax. The presented approach offers the first method that both achieves high-quality video stabilization and is practical enough for consumer applications.
Feng Liu 0015, Michael Gleicher, Jue Wang 0001, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.5
2011 Color compatibility from large datasets
abstract
This paper studies color compatibility theories using large datasets, and develops new tools for choosing colors. There are three parts to this work. First, using on-line datasets, we test new and existing theories of human color preferences. For example, we test whether certain hues or hue templates may be preferred by viewers. Second, we learn quantitative models that score the quality of a five-color set of colors, called a color theme . Such models can be used to rate the quality of a new color theme. Third, we demonstrate simple proto-types that apply a learned model to tasks in color design, including improving existing themes and extracting themes from images.
Peter O'Donovan, Aseem Agarwala, Aaron Hertzmann
ACM Trans. Graph.2
2010 Computational rephotography
abstract
Rephotographers aim to recapture an existing photograph from the same viewpoint. A historical photograph paired with a well-aligned modern rephotograph can serve as a remarkable visualization of the passage of time. However, the task of rephotography is tedious and often imprecise, because reproducing the viewpoint of the original photograph is challenging. The rephotographer must disambiguate between the six degrees of freedom of 3D translation and rotation, and the confounding similarity between the effects of camera zoom and dolly. We present a real-time estimation and visualization technique for rephotography that helps users reach a desired viewpoint during capture. The input to our technique is a reference image taken from the desired viewpoint. The user moves through the scene with a camera and follows our visualization to reach the desired viewpoint. We employ computer vision techniques to compute the relative viewpoint difference. We guide 3D movement using two 2D arrows. We demonstrate the success of our technique by rephotographing historical images and conducting user studies.
Soonmin Bae, Aseem Agarwala, Frédo Durand
ACM Trans. Graph.2
2010 Image warps for artistic perspective manipulation
abstract
Painters and illustrators commonly sketch vanishing points and lines to guide the construction of perspective images. We present a tool that gives users the ability to manipulate perspective in photographs using image space controls similar to those used by artists. Our approach computes a 2D warp guided by constraints based on projective geometry. A user annotates an image by marking a number of image space constraints including planar regions of the scene, straight lines, and associated vanishing points. The user can then use the lines, vanishing points, and other point constraints as handles to control the warp. Our system optimizes the warp such that straight lines remain straight, planar regions transform according to a homography, and the entire mapping is as shape-preserving as possible. While the result of this warp is not necessarily an accurate perspective projection of the scene, it is often visually plausible. We demonstrate how this approach can be used to produce a variety of effects, such as changing the perspective composition of a scene, exploring artistic perspectives not realizable with a camera, and matching perspectives of objects from different images so that they appear consistent for compositing.
Robert Carroll, Aseem Agarwala, Maneesh Agrawala
ACM Trans. Graph.2
2009 Parallax photography: creating 3D cinematic effects from stills
Ke Colin Zheng, Alex Colburn, Aseem Agarwala, Maneesh Agrawala, David Salesin, Brian Curless, Michael F. Cohen
Graphics Interface3
2009 Light field video stabilization
abstract
We describe a method for producing a smooth, stabilized video from the shaky input of a hand-held light field video camera—specifically, a small camera array. Traditional stabilization techniques dampen shake with 2D warps, and thus have limited ability to stabilize a significantly shaky camera motion through a 3D scene. Other recent stabilization techniques synthesize novel views as they would have been seen along a virtual, smooth 3D camera path, but are limited to static scenes. We show that video camera arrays enable much more powerful video stabilization, since they allow changes in viewpoint for a single time instant. Furthermore, we point out that the straightforward approach to light field video stabilization requires computing structure-from-motion, which can be brittle for typical consumer-level video of general dynamic scenes. We present a more robust approach that avoids input camera path reconstruction. Instead, we employ a spacetime optimization that directly computes a sequence of relative poses between the virtual camera and the camera array, while minimizing acceleration of salient visual features in the virtual image plane. We validate our novel method by comparing it to state-of-the-art stabilization software, such as Apple iMovie and 2d3 SteadyMove Pro, on a number of challenging scenes.
Brandon M. Smith 0001, Li Zhang 0003, Hailin Jin, Aseem Agarwala
ICCV4
2009 Optimizing content-preserving projections for wide-angle images
abstract
Any projection of a 3D scene into a wide-angle image unavoidably results in distortion. Current projection methods either bend straight lines in the scene, or locally distort the shapes of scene objects. We present a method that minimizes this distortion by adapting the projection to content in the scene, such as salient scene regions and lines, in order to preserve their shape. Our optimization technique computes a spatially-varying projection that respects user-specified constraints while minimizing a set of energy terms that measure wide-angle image distortion. We demonstrate the effectiveness of our approach by showing results on a variety of wide-angle photographs, as well as comparisons to standard projections.
Robert Carroll, Maneesh Agrawala, Aseem Agarwala
ACM Trans. Graph.3
2009 Content-preserving warps for 3D video stabilization
abstract
We describe a technique that transforms a video from a hand-held video camera so that it appears as if it were taken with a directed camera motion. Our method adjusts the video to appear as if it were taken from nearby viewpoints, allowing 3D camera movements to be simulated. By aiming only for perceptual plausibility, rather than accurate reconstruction, we are able to develop algorithms that can effectively recreate dynamic scenes from a single source video. Our technique first recovers the original 3D camera motion and a sparse set of 3D, static scene points using an off-the-shelf structure-from-motion system. Then, a desired camera path is computed either automatically (e.g., by fitting a linear or quadratic path) or interactively. Finally, our technique performs a least-squares optimization that computes a spatially-varying warp from each input video frame into an output frame. The warp is computed to both follow the sparse displacements suggested by the recovered 3D structure, and avoid deforming the content in the video frame. Our experiments on stabilizing challenging videos of dynamic scenes demonstrate the effectiveness of our technique.
Feng Liu 0015, Michael Gleicher, Hailin Jin, Aseem Agarwala
ACM Trans. Graph.4
2008 Priors for Large Photo Collections and What They Reveal about Cameras
Sujit Kuthirummal, Aseem Agarwala, Dan B. Goldman, Shree K. Nayar
ECCV (4)2
2008 ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments
abstract
Abstract One of the most common tasks in image and video editing is the local adjustment of various properties (e.g., saturation or brightness) of regions within an image or video. Edge‐aware interpolation of user‐drawn scribbles offers a less effort‐intensive approach to this problem than traditional region selection and matting. However, the technique suffers a number of limitations, such as reduced performance in the presence of texture contrast, and the inability to handle fragmented appearances. We significantly improve the performance of edge‐aware interpolation for this problem by adding a boosting‐based classification step that learns to discriminate between the appearance of scribbled pixels. We show that this novel data term in combination with an existing edge‐aware optimization technique achieves substantially better results for the local image and video adjustment problem than edge‐aware interpolation techniques without classification, or related methods such as matting techniques or graph cut segmentation.
Yuanzhen Li, Edward H. Adelson, Aseem Agarwala
Comput. Graph. Forum3
2008 A Comparative Study of Energy Minimization Methods for Markov Random Fields with Smoothness-Based Priors
abstract
Among the most exciting advances in early vision has been the development of efficient energy minimization algorithms for pixel-labeling tasks such as depth or texture computation. It has been known for decades that such problems can be elegantly expressed as Markov random fields, yet the resulting energy minimization problems have been widely viewed as intractable. Recently, algorithms such as graph cuts and loopy belief propagation (LBP) have proven to be very powerful: for example, such methods form the basis for almost all the top-performing stereo methods. However, the tradeoffs among different energy minimization algorithms are still not well understood. In this paper we describe a set of energy minimization benchmarks and use them to compare the solution quality and running time of several common energy minimization algorithms. We investigate three promising recent methods graph cuts, LBP, and tree-reweighted message passing in addition to the well-known older iterated conditional modes (ICM) algorithm. Our benchmark problems are drawn from published energy functions used for stereo, image stitching, interactive segmentation, and denoising. We also provide a general-purpose software interface that allows vision researchers to easily switch between optimization methods. Benchmarks, code, images, and results are available at http://vision.middlebury.edu/MRF/.
Richard Szeliski, Ramin Zabih, Daniel Scharstein, Olga Veksler, Vladimir Kolmogorov, Aseem Agarwala, Marshall F. Tappen, Carsten Rother
IEEE Trans. Pattern Anal. Mach. Intell.6
2008 High-quality motion deblurring from a single image
abstract
We present a new algorithm for removing motion blur from a single image. Our method computes a deblurred image using a unified probabilistic model of both blur kernel estimation and unblurred image restoration. We present an analysis of the causes of common artifacts found in current deblurring methods, and then introduce several novel terms within this probabilistic model that are inspired by our analysis. These terms include a model of the spatial randomness of noise in the blurred image, as well a new local smoothness prior that reduces ringing artifacts by constraining contrast in the unblurred image wherever the blurred image exhibits low contrast. Finally, we describe an effficient optimization scheme that alternates between blur kernel estimation and unblurred image restoration until convergence. As a result of these steps, we are able to produce high quality deblurred results in low computation time. We are even able to produce results of comparable quality to techniques that require additional input images beyond a single blurry photograph, and to methods that require additional hardware.
Qi Shan, Jiaya Jia, Aseem Agarwala
ACM Trans. Graph.3
2007 Using Photographs to Enhance Videos of a Static Scene
Pravin Bhat, C. Lawrence Zitnick, Noah Snavely, Aseem Agarwala, Maneesh Agrawala, Michael F. Cohen, Brian Curless, Sing Bing Kang
Rendering Techniques4
2007 Efficient gradient-domain compositing using quadtrees
abstract
We describe a hierarchical approach to improving the efficiency of gradient-domain compositing , a technique that constructs seamless composites by combining the gradients of images into a vector field that is then integrated to form a composite. While gradient-domain compositing is powerful and widely used, it suffers from poor scalability. Computing an n pixel composite requires solving a linear system with n variables; solving such a large system quickly overwhelms the main memory of a standard computer when performed for multi-megapixel composites, which are common in practice. In this paper we show how to perform gradient-domain compositing approximately by solving an O(p) linear system, where p is the total length of the seams between image regions in the composite; for typical cases, p is O (√ n ). We achieve this reduction by transforming the problem into a space where much of the solution is smooth, and then utilize the pattern of this smoothness to adaptively subdivide the problem domain using quadtrees. We demonstrate the merits of our approach by performing panoramic stitching and image region copy-and-paste in significantly reduced time and memory while achieving visually identical results.
Aseem Agarwala
ACM Trans. Graph.1
2006 Piecewise Image Registration in the Presence of Multiple Large Motions
abstract
We present a technique for computing a dense pixel correspondence between two images of a scene containing multiple large, rigid motions. We model each motion with either a homography (for planar objects) or a fundamental matrix. The various motions in the scene are first extracted by clustering an initial sparse set of correspondences between feature points; we then perform a multi-label graph cut optimization which assigns each pixel to an independent motion and computes its disparity with respect to that motion. We demonstrate our technique on several example scenes and compare our results with previous approaches.
Pravin Bhat, Ke Colin Zheng, Noah Snavely, Aseem Agarwala, Maneesh Agrawala, Michael F. Cohen, Brian Curless
CVPR (2)4
2006 A Comparative Study of Energy Minimization Methods for Markov Random Fields
Richard Szeliski, Ramin Zabih, Daniel Scharstein, Olga Veksler, Vladimir Kolmogorov, Aseem Agarwala, Marshall F. Tappen, Carsten Rother
ECCV (2)6
2006 Photographing long scenes with multi-viewpoint panoramas
abstract
We present a system for producing multi-viewpoint panoramas of long, roughly planar scenes, such as the facades of buildings along a city street, from a relatively sparse set of photographs captured with a handheld still camera that is moved along the scene. Our work is a significant departure from previous methods for creating multi-viewpoint panoramas, which composite thin vertical strips from a video sequence captured by a translating video camera, in that the resulting panoramas are composed of relatively large regions of ordinary perspective. In our system, the only user input required beyond capturing the photographs themselves is to identify the dominant plane of the photographed scene; our system then computes a panorama automatically using Markov Random Field optimization. Users may exert additional control over the appearance of the result by drawing rough strokes that indicate various high-level goals. We demonstrate the results of our system on several scenes, including urban streets, a river bank, and a grocery store aisle.
Aseem Agarwala, Maneesh Agrawala, Michael F. Cohen, David Salesin, Richard Szeliski
ACM Trans. Graph.1
2005 Panoramic video textures
abstract
This paper describes a mostly automatic method for taking the output of a single panning video camera and creating a panoramic video texture (PVT): a video that has been stitched into a single, wide field of view and that appears to play continuously and indefinitely. The key problem in creating a PVT is that although only a portion of the scene has been imaged at any given time, the output must simultaneously portray motion throughout the scene. Like previous work in video textures, our method employs min-cut optimization to select fragments of video that can be stitched together both spatially and temporally. However, it differs from earlier work in that the optimization must take place over a much larger set of data. Thus, to create PVTs, we introduce a dynamic programming step, followed by a novel hierarchical min-cut optimization algorithm. We also use gradient-domain compositing to further smooth boundaries between video fragments. We demonstrate our results with an interactive viewer in which users can interactively pan and zoom on high-resolution PVTs.
Aseem Agarwala, Ke Colin Zheng, Christopher Joseph Pal, Maneesh Agrawala, Michael F. Cohen, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.1
2004 Interactive digital photomontage
abstract
We describe an interactive, computer-assisted framework for combining parts of a set of photographs into a single composite picture, a process we call "digital photomontage." Our framework makes use of two techniques primarily: graph-cut optimization, to choose good seams within the constituent images so that they can be combined as seamlessly as possible; and gradient-domain fusion, a process based on Poisson equations, to further reduce any remaining visible artifacts in the composite. Also central to the framework is a suite of interactive tools that allow the user to specify a variety of high-level image objectives, either globally across the image, or locally through a painting-style interface. Image objectives are applied independently at each pixel location and generally involve a function of the pixel values (such as "maximum contrast") drawn from that same location in the set of source images. Typically, a user applies a series of image objectives iteratively in order to create a finished composite. The power of this framework lies in its generality; we show how it can be used for a wide variety of applications, including "selective composites" (for instance, group photos in which everyone looks their best), relighting, extended depth of field, panoramic stitching, clean-plate production, stroboscopic visualization of movement, and time-lapse mosaics.
Aseem Agarwala, Mira Dontcheva, Maneesh Agrawala, Steven Mark Drucker, Alex Colburn, Brian Curless, David Salesin, Michael F. Cohen
ACM Trans. Graph.1
2004 Keyframe-based tracking for rotoscoping and animation
abstract
We describe a new approach to rotoscoping --- the process of tracking contours in a video sequence --- that combines computer vision with user interaction. In order to track contours in video, the user specifies curves in two or more frames; these curves are used as keyframes by a computer-vision-based tracking algorithm. The user may interactively refine the curves and then restart the tracking algorithm. Combining computer vision with user interaction allows our system to track any sequence with significantly less effort than interpolation-based systems --- and with better reliability than "pure" computer vision systems. Our tracking algorithm is cast as a spacetime optimization problem that solves for time-varying curve shapes based on an input video sequence and user-specified constraints. We demonstrate our system with several rotoscoped examples. Additionally, we show how these rotoscoped contours can be used to help create cartoon animation by attaching user-drawn strokes to the tracked contours.
Aseem Agarwala, Aaron Hertzmann, David Salesin, Steven M. Seitz
ACM Trans. Graph.1
2002 Video matting of complex scenes
abstract
This paper describes a new framework for video matting, the process of pulling a high-quality alpha matte and foreground from a video sequence. The framework builds upon techniques in natural image matting, optical flow computation, and background estimation. User interaction is comprised of garbage matte specification if background estimation is needed, and hand-drawn keyframe segmentations into "foreground," "background" and "unknown". The segmentations, called trimaps, are interpolated across the video volume using forward and backward optical flow. Competing flow estimates are combined based on information about where flow is likely to be accurate. A Bayesian matting technique uses the flowed trimaps to yield high-quality mattes of moving foreground elements with complex boundaries filmed by a moving camera. A novel technique for smoke matte extraction is also demonstrated.
Yung-Yu Chuang, Aseem Agarwala, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.2
2000 Tangible interaction + graphical interpretation: a new approach to 3D modeling
abstract
Construction toys are a superb medium for geometric models. We argue that such toys, suitably instrumented or sensed, could be the inspiration for a new generation of easy-to-use, tangible modeling systems—especially if the tangible modeling is combined with graphical-interpretation techniques for enhancing nascent models automatically. The three key technologies needed to realize this idea are embedded computation, vision-based acquisition, and graphical interpretation. We sample these technologies in the context of two novel modeling systems: physical building blocks that self-describe, interpret, and decorate the structures into which they are assembled; and a system for scanning, interpreting, and animating clay figures.
David B. Anderson, James L. Frankel, Joe Marks, Aseem Agarwala, Paul A. Beardsley, Jessica K. Hodgins, Darren Leigh, Kathy Ryall, Eddie Sullivan, Jonathan S. Yedidia
SIGGRAPH4