Adam Finkelstein

dblp:f/AdamFinkelstein · DBLP profile ↗
← Back
76ranked-venue papers
2as first author
11since 2021 · last 2026
0000-0001-9422-5363ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 66 · 2 first-author · 11 since 2021Human-computer interaction and ubiquitous computing · 20 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 1 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Lucky High Dynamic Range Smartphone Imaging
Baiang Li, Ruyu Yan, Ethan Tseng, Zhoutong Zhang, Adam Finkelstein, Jiawen Chen 0001, Felix Heide
ACM Trans. Graph.5
2025 Temporally Smooth Mesh Extraction for Procedural Scenes with Long-Range Camera Trajectories using Spacetime Octrees
abstract
The procedural occupancy function is a flexible and compact representation for creating 3D scenes. For rasterization and other tasks, it is often necessary to extract a mesh that represents the shape. Unbounded scenes with long-range camera trajectories, such as flying through a forest, pose a unique challenge for mesh extraction. A single static mesh representing all the geometric detail necessary for the full camera path can be prohibitively large. Therefore, independent meshes can be extracted for different camera views, but this approach may lead to popping artifacts during transitions. We propose a temporally coherent method for extracting meshes suitable for long-range camera trajectories in unbounded scenes represented by an occupancy function. The key idea is to perform 4D mesh extraction using a new spacetime tree structure called a binary-octree. Experiments show that, compared to existing baseline methods, our method offers superior visual consistency at a comparable cost. The code and the supplementary video for this paper are available at https://github.com/princeton-vl/BinocMesher.
Zeyu Ma 0004, Adam Finkelstein, Jia Deng 0001
SIGGRAPH Asia2
2024 Corn: Co-Trained Full- and No-Reference Speech Quality Assessment
abstract
Perceptual evaluation constitutes a crucial aspect of various audio-processing tasks. Full reference (FR) or similarity-based metrics rely on high-quality reference recordings, to which lower-quality or corrupted versions of the recording may be compared for evaluation. In contrast, no-reference (NR) metrics evaluate a recording without relying on a reference. Both the FR and NR approaches exhibit advantages and drawbacks relative to each other. In this paper, we present a novel framework called CORN that amalgamates these dual approaches, concurrently training both FR and NR models together. After training, the models can be applied independently. We evaluate CORN by predicting several common objective metrics and across two different architectures. The NR model trained using CORN has access to a reference recording during training, and thus, as one would expect, it consistently outperforms baseline NR models trained independently. Perhaps even more remarkable is that the CORN FR model also outperforms its baseline counterpart, even though it relies on the same training data and the same model architecture. Thus, a single training regime produces two independently useful models, each outperforming independently trained models.
Pranay Manocha, Donald Williamson, Adam Finkelstein
ICASSP3
2024 GR0: Self-Supervised Global Representation Learning for Zero-Shot Voice Conversion
abstract
Research in generative self-supervised learning (SSL) has largely focused on local embeddings for tokenized sequences. We introduce a generative SSL framework that learns a global representation that is disentangled from local embeddings. We apply this technique to jointly learn a global speaker embedding and a zero-shot voice converter. The converter modifies recorded speech to sound as if it were spoken by a different person while preserving the content, using only a short reference clip unavailable to the model during training. Listening experiments conducted on an unseen dataset show that our models significantly outperform SOTA baselines in both quality and speaker similarity for various datasets and unseen languages.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP3
2022 SQAPP: No-Reference Speech Quality Assessment Via Pairwise Preference
abstract
Automatic speech quality assessment remains challenging, as we lack complete models of human auditory perception. Many existing full-reference models correlate well with human perception, but cannot be used in real-world scenarios where ground truth clean reference recordings are not available. On the other hand no-reference metrics typically suffer from several shortcomings, such as lack of robustness to unseen perturbations and reliance on (limited) labeled data for training. Moreover, noise or large variance among the labels makes it difficult to learn generalizable representations, especially for recordings with subtle differences. This paper proposes a learning framework for estimating the quality of a recording without any reference, and without any human judgments. The main component of this framework is a pairwise quality-preference strategy that reduces label noise, thereby making learning more robust. From pairwise preferences, we first learn a content invariant quality ordering; and then we re-target the model to predict quality on an absolute scale. We show that the resulting learned metric is well-calibrated with human judgments. Since it is a deep network, the metric is differentiable, making it suitable as a loss function for downstream tasks. For example, we show that adding this metric to an existing speech enhancement method yields significant improvement.
Pranay Manocha, Zeyu Jin, Adam Finkelstein
ICASSP3
2022 Controllable Speech Representation Learning Via Voice Conversion and AIC Loss
abstract
Speech representation learning transforms speech into features that are suitable for downstream tasks, e.g. speech recognition, phoneme classification, or speaker identification. For such recognition tasks, a representation can be lossy (non-invertible), which is typical of BERT-like self-supervised models. However, when used for synthesis tasks, we find these lossy representations prove to be insufficient to plausibly reconstruct the input signal. This paper introduces a method for invertible and controllable speech representation learning based on disentanglement. The representation can be decoded into a signal perceptually identical to the original. Moreover, its disentangled components (content, pitch, speaker identity, and energy) can be controlled independently to alter the synthesis result. Our model builds upon a zero-shot voice conversion model AutoVC-F0, in which we introduce alteration invariant content loss (AIC loss) and adversarial training (GAN). Through objective measures and subjective tests, we show that our formulation offers significant improvement in voice conversion sound quality as well as more precise control over the disentangled features.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP3
2022 Audio Similarity is Unreliable as a Proxy for Audio Quality
abstract
Many audio processing tasks require perceptual assessment.However, the time and expense of obtaining "gold standard" human judgments limit the availability of such data.Most applications incorporate full reference or other similarity-based metrics (e.g.PESQ) that depend on a clean reference.Researchers have relied on such metrics to evaluate and compare various proposed methods, often concluding that small, measured differences imply one is more effective than another.This paper demonstrates several practical scenarios where similarity metrics fail to agree with human perception, because they: (1) vary with clean references; (2) rely on attributes that humans factor out when considering quality, and (3) are sensitive to imperceptible signal level differences.In those scenarios, we show that no-reference metrics do not suffer from such shortcomings and correlate better with human perception.We conclude therefore that similarity serves as an unreliable proxy for audio quality.
Pranay Manocha, Zeyu Jin, Adam Finkelstein
INTERSPEECH3
2022 Learning from Shader Program Traces
abstract
Abstract Deep learning for image processing typically treats input imagery as pixels in some color space. This paper proposes instead to learn from program traces of procedural fragment shaders – programs that generate images. At each pixel, we collect the intermediate values computed at program execution, and these data form the input to the learned model. We investigate this learning task for a variety of applications: our model can learn to predict a low‐noise output image from shader programs that exhibit sampling noise; this model can also learn from a simplified shader program that approximates the reference solution with less computation, as well as learn the output of postprocessing filters like defocus blur and edge‐aware sharpening. Finally we show that the idea of learning from program traces can even be applied to non‐imagery simulations of flocks of boids. Our experiments on a variety of shaders show quantitatively and qualitatively that models learned from program traces outperform baseline models learned from RGB color augmented with hand‐picked shader‐specific featues like normals, depth, and diffuse and specular color. We also conduct a series of analyses that show certain features within the trace are more important, and even learning from a small subset of the trace outperforms the baselines.
Yuting Yang 0004, Connelly Barnes, Adam Finkelstein
Comput. Graph. Forum3
2022 Aδ: autodiff for discontinuous programs - applied to shaders
abstract
Over the last decade, automatic differentiation (AD) has profoundly impacted graphics and vision applications --- both broadly via deep learning and specifically for inverse rendering. Traditional AD methods ignore gradients at discontinuities, instead treating functions as continuous. Rendering algorithms intrinsically rely on discontinuities, crucial at object silhouettes and in general for any branching operation. Researchers have proposed fully- automatic differentiation approaches for handling discontinuities by restricting to affine functions, or semi- automatic processes restricted either to invertible functions or to specialized applications like vector graphics. This paper describes a compiler-based approach to extend reverse mode AD so as to accept arbitrary programs involving discontinuities. Our novel gradient rules generalize differentiation to work correctly, assuming there is a single discontinuity in a local neighborhood, by approximating the prefiltered gradient over a box kernel oriented along a 1D sampling axis. We describe when such approximation rules are first-order correct, and show that this correctness criterion applies to a relatively broad class of functions. Moreover, we show that the method is effective in practice for arbitrary programs, including features for which we cannot prove correctness. We evaluate this approach on procedural shader programs, where the task is to optimize unknown parameters in order to match a target image, and our method outperforms baselines in terms of both convergence and efficiency. Our compiler outputs gradient programs in TensorFlow, PyTorch (for quick prototypes) and Halide with an optional auto-scheduler (for efficiency). The compiler also outputs GLSL that renders the target image, allowing users to interactively modify and animate the shader, which would otherwise be cumbersome in other representations such as triangle meshes or vector art.
Yuting Yang 0004, Connelly Barnes, Andrew Adams, Adam Finkelstein
ACM Trans. Graph.4
2021 CDPAM: Contrastive Learning for Perceptual Audio Similarity
abstract
Many speech processing methods based on deep learning require an automatic and differentiable audio metric for the loss function. The DPAM approach of Manocha et al. [1] learns a full-reference metric trained directly on human judgments, and thus correlates well with human perception. However, it requires a large number of human annotations and does not generalize well outside the range of perturbations on which it was trained. This paper introduces CDPAM –a metric that builds on and advances DPAM. The primary improvement is to combine contrastive learning and multi-dimensional representations to build robust models from limited data. In addition, we collect human judgments on triplet comparisons to improve generalization to a broader range of audio perturbations. CDPAM correlates well with human responses across nine varied datasets. We also show that adding this metric to existing speech synthesis and enhancement methods yields significant improvement, as measured by objective and subjective tests.
Pranay Manocha, Zeyu Jin, Richard Zhang 0001, Adam Finkelstein
ICASSP4
2021 Bandwidth Extension is All You Need
abstract
Speech generation and enhancement have seen recent breakthroughs in quality thanks to deep learning. These methods typically operate at a limited sampling rate of 16-22kHz due to computational complexity and available datasets. This limitation imposes a gap between the output of such methods and that of high-fidelity (≥44kHz) real-world audio applications. This paper proposes a new bandwidth extension (BWE) method that expands 8-16kHz speech signals to 48kHz. The method is based on a feed-forward WaveNet architecture trained with a GAN-based deep feature loss. A mean-opinion-score (MOS) experiment shows significant improvement in quality over state-of-the-art BWE methods. An AB test reveals that our 16-to-48kHz BWE is able to achieve fidelity that is typically indistinguishable from real high-fidelity recordings. We use our method to enhance the output of recent speech generation and denoising methods, and experiments demonstrate significant improvement in sound quality over these baselines. We propose this as a general approach to narrow the gap between generated speech and recorded speech, without the need to adapt such methods to higher sampling rates.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP3
2020 Acoustic Matching By Embedding Impulse Responses
abstract
The goal of acoustic matching is to transform an audio recording made in one acoustic environment to sound as if it had been recorded in a different environment, based on reference audio from the target environment. This paper introduces a deep learning solution for two parts of the acoustic matching problem. First, we characterize acoustic environments by mapping audio into a low-dimensional embedding invariant to speech content and speaker identity. Next, a waveform-to-waveform neural network conditioned on this embedding learns to transform an input waveform to match the acoustic qualities encoded in the target embedding. Listening tests on both simulated and real environments show that the proposed approach improves on state-of-the-art baseline methods.
Jiaqi Su, Zeyu Jin, Adam Finkelstein
ICASSP3
2020 A Differentiable Perceptual Audio Metric Learned from Just Noticeable Differences
abstract
Many audio processing tasks require perceptual assessment.The "gold standard" of obtaining human judgments is timeconsuming, expensive, and cannot be used as an optimization criterion.On the other hand, automated metrics are efficient to compute but often correlate poorly with human judgment, particularly for audio differences at the threshold of human detection.In this work, we construct a metric by fitting a deep neural network to a new large dataset of crowdsourced human judgments.Subjects are prompted to answer a straightforward, objective question: are two recordings identical or not?These pairs are algorithmically generated under a variety of perturbations, including noise, reverb, and compression artifacts; the perturbation space is probed with the goal of efficiently identifying the just-noticeable difference (JND) level of the subject.We show that the resulting learned metric is well-calibrated with human judgments, outperforming baseline methods.Since it is a deep network, the metric is differentiable, making it suitable as a loss function for other tasks.Thus, simply replacing an existing loss (e.g., deep feature loss) with our metric yields significant improvement in a denoising network, as measured by subjective pairwise comparison.
Pranay Manocha, Adam Finkelstein, Richard Zhang 0001, Nicholas J. Bryan, Gautham J. Mysore, Zeyu Jin
INTERSPEECH2
2020 HiFi-GAN: High-Fidelity Denoising and Dereverberation Based on Speech Deep Features in Adversarial Networks
abstract
Real-world audio recordings are often degraded by factors such as noise, reverberation, and equalization distortion.This paper introduces HiFi-GAN, a deep learning method to transform recorded speech to sound as though it had been recorded in a studio.We use an end-to-end feed-forward WaveNet architecture, trained with multi-scale adversarial discriminators in both the time domain and the time-frequency domain.It relies on the deep feature matching losses of the discriminators to improve the perceptual quality of enhanced speech.The proposed model generalizes well to new speakers, new speech content, and new environments.It significantly outperforms state-of-the-art baseline methods in both objective and subjective experiments.
Jiaqi Su, Zeyu Jin, Adam Finkelstein
INTERSPEECH3
2020 Pose2Pose: pose selection and transfer for 2D character animation
abstract
An artist faces two challenges when creating a 2D animated character to mimic a specific human performance. First, the artist must design and draw a collection of artwork depicting portions of the character in a suitable set of poses, for example arm and hand poses that can be selected and combined to express the range of gestures typical for that person. Next, to depict a specific performance, the artist must select and position the appropriate set of artwork at each moment of the animation. This paper presents a system that addresses these challenges by leveraging video of the target human performer. Our system tracks arm and hand poses in an example video of the target. The UI displays clusters of these poses to help artists select representative poses that capture the actor's style and personality. From this mapping of pose data to character artwork, our system can generate an animation from a new performance video. It relies on a dynamic programming algorithm to optimize for smooth animations that match the poses found in the video. Artists used our system to create four 2D characters and were pleased with the final automatically animated results. We also describe additional applications addressing audio-driven or text-based animations.
Nora S. Willett, Hijung Shin, Zeyu Jin, Wilmot Li, Adam Finkelstein
IUI5
2019 Learning Bandwidth Expansion Using Perceptually-motivated Loss
abstract
We introduce a perceptually motivated approach to bandwidth expansion for speech. Our method pairs a new 3-way split variant of the FFTNet neural vocoder structure with a perceptual loss function, combining objectives from both the time and frequency domains. Mean opinion score tests show that it outperforms baseline methods from both domains, even for extreme bandwidth expansion.
Berthy Feng, Zeyu Jin, Jiaqi Su, Adam Finkelstein
ICASSP4
2019 Perceptually-motivated Environment-specific Speech Enhancement
abstract
This paper introduces a deep learning approach to enhance speech recordings made in a specific environment. A single neural network learns to ameliorate several types of recording artifacts, including noise, reverberation, and non-linear equalization. The method relies on a new perceptual loss function that combines adversarial loss with spectrogram features. Both subjective and objective evaluations show that the proposed approach improves on state-of-the-art baseline methods.
Jiaqi Su, Adam Finkelstein, Zeyu Jin
ICASSP2
2019 High-Precision Localization Using Ground Texture
abstract
Location-aware applications play an increasingly critical role in everyday life. However, satellite-based localization (e.g., GPS) has limited accuracy and can be unusable in dense urban areas and indoors. We introduce an image-based global localization system that is accurate to a few millimeters and performs reliable localization both indoors and outside. The key idea is to capture and index distinctive local keypoints in ground textures. This is based on the observation that ground textures including wood, carpet, tile, concrete, and asphalt may look random and homogeneous, but all contain cracks, scratches, or unique arrangements of fibers. These imperfections are persistent, and can serve as local features. Our system incorporates a downward-facing camera to capture the fine texture of the ground, together with an image processing pipeline that locates the captured texture patch in a compact database constructed offline. We demonstrate the capability of our system to robustly, accurately, and quickly locate test images on various types of outdoor and indoor ground surfaces. This paper contains a supplementary video. All datasets and code are available online at microgps.cs.princeton.edu.
Linguang Zhang, Adam Finkelstein, Szymon Rusinkiewicz
ICRA2
2019 Text-based editing of talking-head video
abstract
Editing talking-head video to change the speech content or to remove filler words is challenging. We propose a novel method to edit talking-head video based on its transcript to produce a realistic output video in which the dialogue of the speaker has been modified, while maintaining a seamless audio-visual flow (i.e. no jump cuts). Our method automatically annotates an input talking-head video with phonemes, visemes, 3D face pose and geometry, reflectance, expression and scene illumination per frame. To edit a video, the user has to only edit the transcript, and an optimization strategy then chooses segments of the input corpus as base material. The annotated parameters corresponding to the selected segments are seamlessly stitched together and used to produce an intermediate video representation in which the lower half of the face is rendered with a parametric face model. Finally, a recurrent video generation network transforms this representation to a photorealistic video that matches the edited transcript. We demonstrate a large variety of edits, such as the addition, removal, and alteration of words, as well as convincing language translation and full sentence synthesis.
Ohad Fried, Ayush Tewari, Michael Zollhöfer, Adam Finkelstein, Eli Shechtman, Dan B. Goldman, Kyle Genova, Zeyu Jin, Christian Theobalt, Maneesh Agrawala
ACM Trans. Graph.4
2018 PairedCycleGAN: Asymmetric Style Transfer for Applying and Removing Makeup
abstract
This paper introduces an automatic method for editing a portrait photo so that the subject appears to be wearing makeup in the style of another person in a reference photo. Our unsupervised learning approach relies on a new framework of cycle-consistent generative adversarial networks. Different from the image domain transfer problem, our style transfer problem involves two asymmetric functions: a forward function encodes example-based style transfer, whereas a backward function removes the style. We construct two coupled networks to implement these functions - one that transfers makeup style and a second that can remove makeup - such that the output of their successive application to an input photo will match the input. The learned style network can then quickly apply an arbitrary makeup style to an arbitrary photo. We demonstrate the effectiveness on a broad range of portraits and styles.
Huiwen Chang, Jingwan Lu, Fisher Yu 0001, Adam Finkelstein
CVPR4
2018 Fftnet: A Real-Time Speaker-Dependent Neural Vocoder
abstract
We introduce FFTNet, a deep learning approach synthesizing audio waveforms. Our approach builds on the recent WaveNet project, which showed that it was possible to synthesize a natural sounding audio waveform directly from a deep convolutional neural network. FFTNet offers two improvements over WaveNet. First it is substantially faster, allowing for real-time synthesis of audio waveforms. Second, when used as a vocoder, the resulting speech sounds more natural, as measured via a “mean opinion score” test.
Zeyu Jin, Adam Finkelstein, Gautham J. Mysore, Jingwan Lu
ICASSP2
2018 A Mixed-Initiative Interface for Animating Static Pictures
abstract
We present an interactive tool to animate the visual elements of a static picture, based on simple sketch-based markup. While animated images enhance websites, infographics, logos, e-books, and social media, creating such animations from still pictures is difficult for novices and tedious for experts. Creating automatic tools is challenging due to ambiguities in object segmentation, relative depth ordering, and non-existent temporal information. With a few user drawn scribbles as input, our mixed initiative creative interface extracts repetitive texture elements in an image, and supports animating them. Our system also facilitates the creation of multiple layers to enhance depth cues in the animation. Finally, after analyzing the artwork during segmentation, several animation processes automatically generate kinetic textures that are spatio-temporally coherent with the source image. Our results, as well as feedback from our user evaluation, suggest that our system effectively allows illustrators and animators to add life to still images in a broad range of visual styles.
Nora S. Willett, Rubaiat Habib Kazi, George W. Fitzmaurice, Adam Finkelstein, Tovi Grossman
UIST5
2017 Secondary Motion for Performed 2D Animation
abstract
When bringing animated characters to life, artists often augment the primary motion of a figure by adding secondary animation -- subtle movement of parts like hair, foliage or cloth that complements and emphasizes the primary motion. Traditionally, artists add secondary motion to animated illustrations only through arduous manual effort, and often eschew it entirely. Emerging ``live' performance applications allow both novices and experts to perform the primary motion of a character, but only a virtuoso performer could manage the degrees of freedom needed to specify both primary and secondary motion together. This paper introduces physically-inspired rigs that propagate the primary motion of layered, illustrated characters to produce plausible secondary motion. These composable elements are rigged and controlled via a small number of parameters to produce an expressive range of effects. Our approach supports a variety of the most common secondary effects, which we demonstrate with an assortment of characters of varying complexity.
Nora S. Willett, Wilmot Li, Jovan Popovic, Floraine Berthouzoz, Adam Finkelstein
UIST5
2017 Triggering Artwork Swaps for Live Animation
abstract
Live animation of 2D characters is a new form of storytelling that has started to appear on streaming platforms and broadcast TV. Unlike traditional animation, human performers control characters in real time so that they can respond and improvise to live events. Current live animation systems provide a range of animation controls, such as camera input to drive head movements, audio for lip sync, and keyboard shortcuts to trigger discrete pose changes via artwork swaps. However, managing all of these controls during a live performance is challenging. In this work, we present a new interactive system that specifically addresses the problem of triggering artwork swaps in live settings. Our key contributions are the design of a multi-touch triggering interface that overlays visual triggers around a live preview of the character, and a predictive triggering model that leverages practice performances to suggest pose transitions during live performances. We evaluate our system with quantitative experiments, a user study with novice participants, and interviews with professional animators.
Nora S. Willett, Wilmot Li, Jovan Popovic, Adam Finkelstein
UIST4
2017 VoCo: text-based insertion and replacement in audio narration
abstract
Editing audio narration using conventional software typically involves many painstaking low-level manipulations. Some state of the art systems allow the editor to work in a text transcript of the narration, and perform select, cut, copy and paste operations directly in the transcript; these operations are then automatically applied to the waveform in a straightforward manner. However, an obvious gap in the text-based interface is the ability to type new words not appearing in the transcript, for example inserting a new word for emphasis or replacing a misspoken word. While high-quality voice synthesizers exist today, the challenge is to synthesize the new word in a voice that matches the rest of the narration. This paper presents a system that can synthesize a new word or short phrase such that it blends seamlessly in the context of the existing narration. Our approach is to use a text to speech synthesizer to say the word in a generic voice, and then use voice conversion to convert it into a voice that matches the narration. Offering a range of degrees of control to the editor, our interface supports fully automatic synthesis, selection among a candidate set of alternative pronunciations, fine control over edit placements and pitch profiles, and even guidance by the editors own voice. The paper presents studies showing that the output of our method is preferred over baseline methods and often indistinguishable from the original voice.
Zeyu Jin, Gautham J. Mysore, Stephen DiVerdi, Jingwan Lu, Adam Finkelstein
ACM Trans. Graph.5
2016 Cute: A concatenative method for voice conversion using exemplar-based unit selection
abstract
State-of-the art voice conversion methods re-synthesize voice from spectral representations such as MFCCs and STRAIGHT, thereby introducing muffled artifacts. We propose a method that circumvents this concern using concatenative synthesis coupled with exemplar-based unit selection. Given parallel speech from source and target speakers as well as a new query from the source, our method stitches together pieces of the target voice. It optimizes for three goals: matching the query, using long consecutive segments, and smooth transitions between the segments. To achieve these goals, we perform unit selection at the frame level and introduce triphone-based preselection that greatly reduces computation and enforces selection of long, contiguous pieces. Our experiments show that the proposed method has better quality than baseline methods, while preserving high individuality.
Zeyu Jin, Adam Finkelstein, Stephen DiVerdi, Jingwan Lu, Gautham J. Mysore
ICASSP2
2016 Automatic triage for a photo series
abstract
People often take a series of nearly redundant pictures to capture a moment or scene. However, selecting photos to keep or share from a large collection is a painful chore. To address this problem, we seek a relative quality measure within a series of photos taken of the same scene, which can be used for automatic photo triage. Towards this end, we gather a large dataset comprised of photo series distilled from personal photo albums. The dataset contains 15, 545 unedited photos organized in 5,953 series. By augmenting this dataset with ground truth human preferences among photos within each series, we establish a benchmark for measuring the effectiveness of algorithmic models of how people select photos. We introduce several new approaches for modeling human preference based on machine learning. We also describe applications for the dataset and predictor, including a smart album viewer, automatic photo enhancement, and providing overviews of video clips.
Huiwen Chang, Fisher Yu 0001, Jue Wang 0001, Douglas Ashley, Adam Finkelstein
ACM Trans. Graph.5
2016 Perspective-aware manipulation of portrait photos
abstract
This paper introduces a method to modify the apparent relative pose and distance between camera and subject given a single portrait photo. Our approach fits a full perspective camera and a parametric 3D head model to the portrait, and then builds a 2D warp in the image plane to approximate the effect of a desired change in 3D. We show that this model is capable of correcting objectionable artifacts such as the large noses sometimes seen in "selfies," or to deliberately bring a distant camera closer to the subject. This framework can also be used to re-pose the subject, as well as to create stereo pairs from an input portrait. We show convincing results on both an existing dataset as well as a new dataset we captured to validate our method.
Ohad Fried, Eli Shechtman, Dan B. Goldman, Adam Finkelstein
ACM Trans. Graph.4
2015 Finding distractors in images
abstract
We propose a new computer vision task we call “distractor prediction.” Distractors are the regions of an image that draw attention away from the main subjects and reduce the overall image quality. Removing distractors-for example, using in-painting - can improve the composition of an image. In this work we created two datasets of images with user annotations to identify the characteristics of distractors. We use these datasets to train an algorithm to predict distractor maps. Finally, we use our predictor to automatically enhance images.
Ohad Fried, Eli Shechtman, Dan B. Goldman, Adam Finkelstein
CVPR4
2015 IsoMatch: Creating Informative Grid Layouts
abstract
Abstract Collections of objects such as images are often presented visually in a grid because it is a compact representation that lends itself well for search and exploration. Most grid layouts are sorted using very basic criteria, such as date or filename. In this work we present a method to arrange collections of objects respecting an arbitrary distance measure. Pairwise distances are preserved as much as possible, while still producing the specific target arrangement which may be a 2D grid, the surface of a sphere, a hierarchy, or any other shape. We show that our method can be used for infographics, collection exploration, summarization, data visualization, and even for solving problems such as where to seat family members at a wedding. We present a fast algorithm that can work on large collections and quantitatively evaluate how well distances are preserved.
Ohad Fried, Stephen DiVerdi, Maciej Halber, Elena Sizikova, Adam Finkelstein
Comput. Graph. Forum5
2015 Palette-based photo recoloring
abstract
Image editing applications offer a wide array of tools for color manipulation. Some of these tools are easy to understand but offer a limited range of expressiveness. Other more powerful tools are time consuming for experts and inscrutable to novices. Researchers have described a variety of more sophisticated methods but these are typically not interactive, which is crucial for creative exploration. This paper introduces a simple, intuitive and interactive tool that allows non-experts to recolor an image by editing a color palette. This system is comprised of several components: a GUI that is easy to learn and understand, an efficient algorithm for creating a color palette from an image, and a novel color transfer algorithm that recolors the image based on a user-modified palette. We evaluate our approach via a user study, showing that it is faster and easier to use than two alternatives, and allows untrained users to achieve results comparable to those of experts using professional software.
Huiwen Chang, Ohad Fried, Yiming Liu 0001, Stephen DiVerdi, Adam Finkelstein
ACM Trans. Graph.5
2014 DecoBrush: drawing structured decorative patterns by example
abstract
Structured decorative patterns are common ornamentations in a variety of media like books, web pages, greeting cards and interior design. Creating such art from scratch using conventional software is time consuming for experts and daunting for novices. We introduce DecoBrush, a data-driven drawing system that generalizes the conventional digital "painting" concept beyond the scope of natural media to allow synthesis of structured decorative patterns following user-sketched paths. The user simply selects an example library and draws the overall shape of a pattern. DecoBrush then synthesizes a shape in the style of the exemplars but roughly matching the overall shape. If the designer wishes to alter the result, DecoBrush also supports user-guided refinement via simple drawing and erasing tools. For a variety of example styles, we demonstrate high-quality user-constrained synthesized patterns that visually resemble the exemplars while exhibiting plausible structural variations.
Jingwan Lu, Connelly Barnes, Connie Wan, Paul Asente, Radomír Mech, Adam Finkelstein
ACM Trans. Graph.6
2013 Affect and Creative Performance on Crowdsourcing Platforms
abstract
Performance on crowd sourcing platforms varies greatly, especially for tasks requiring significant cognitive effort or creative insight. Researchers have proposed several techniques to address these problems, yet few have considered the role of affect, despite the well-established link between positive affect and creative performance. In this paper, we examine two affective techniques to boost creativity on crowd sourcing platforms - affective priming and affective pre-screening. Across three experiments, we find divergent results, depending on which technique is used. We find that not all happy crowd workers are alike. Those that are primed to feel happy exhibit enhanced creative performance, whereas those that merely report feeling happy exhibit impaired creative performance. We examine these findings in light of preexisting research on creativity, affect, and mood saliency. Lastly, we show how our findings have implications not only for crowd sourcing platforms, but also for other human-computer interaction scenarios that involve affect and creative performance.
Robert R. Morris, Mira Dontcheva, Adam Finkelstein, Elizabeth Gerber
ACII3
2013 Pixelated image abstraction with integrated user constraints
Timothy Gerstner, Douglas DeCarlo, Marc Alexa, Adam Finkelstein, Yotam I. Gingold, Andrew Nealen
Comput. Graph.4
2013 A no-reference metric for evaluating the quality of motion deblurring
abstract
Methods to undo the effects of motion blur are the subject of intense research, but evaluating and tuning these algorithms has traditionally required either user input or the availability of ground-truth images. We instead develop a metric for automatically predicting the perceptual quality of images produced by state-of-the-art deblurring algorithms. The metric is learned based on a massive user study, incorporates features that capture common deblurring artifacts, and does not require access to the original images (i.e., is "noreference"). We show that it better matches user-supplied rankings than previous approaches to measuring quality, and that in most cases it outperforms conventional full-reference image-similarity measures. We demonstrate applications of this metric to automatic selection of optimal algorithms and parameters, and to generation of fused images that combine multiple deblurring results.
Yiming Liu 0001, Jue Wang 0001, Sunghyun Cho, Adam Finkelstein, Szymon Rusinkiewicz
ACM Trans. Graph.4
2013 RealBrush: painting with examples of physical media
abstract
Conventional digital painting systems rely on procedural rules and physical simulation to render paint strokes. We present an interactive, data-driven painting system that uses scanned images of real natural media to synthesize both new strokes and complex stroke interactions, obviating the need for physical simulation. First, users capture images of real media, including examples of isolated strokes, pairs of overlapping strokes, and smudged strokes. Online, the user inputs a new stroke path, and our system synthesizes its 2D texture appearance with optional smearing or smudging when strokes overlap. We demonstrate high-fidelity paintings that closely resemble the captured media style, and also quantitatively evaluate our synthesis quality via user studies.
Jingwan Lu, Connelly Barnes, Stephen DiVerdi, Adam Finkelstein
ACM Trans. Graph.4
2012 HelpingHand: example-based stroke stylization
abstract
Digital painters commonly use a tablet and stylus to drive software like Adobe Photoshop. A high quality stylus with 6 degrees of freedom (DOFs: 2D position, pressure, 2D tilt, and 1D rotation) coupled to a virtual brush simulation engine allows skilled users to produce expressive strokes in their own style. However, such devices are difficult for novices to control, and many people draw with less expensive (lower DOF) input devices. This paper presents a data-driven approach for synthesizing the 6D hand gesture data for users of low-quality input devices. Offline, we collect a library of strokes with 6D data created by trained artists. Online, given a query stroke as a series of 2D positions, we synthesize the 4D hand pose data at each sample based on samples from the library that locally match the query. This framework optionally can also modify the stroke trajectory to match characteristic shapes in the style of the library. Our algorithm outputs a 6D trajectory that can be fed into any virtual brush stroke engine to make expressive strokes for novices or users of limited hardware.
Jingwan Lu, Fisher Yu 0001, Adam Finkelstein, Stephen DiVerdi
ACM Trans. Graph.3
2011 Perceptual models of viewpoint preference
abstract
The question of what are good views of a 3D object has been addressed by numerous researchers in perception, computer vision, and computer graphics. This has led to a large variety of measures for the goodness of views as well as some special-case viewpoint selection algorithms. In this article, we leverage the results of a large user study to optimize the parameters of a general model for viewpoint goodness, such that the fitted model can predict people's preferred views for a broad range of objects. Our model is represented as a combination of attributes known to be important for view selection, such as projected model area and silhouette length. Moreover, this framework can easily incorporate new attributes in the future, based on the data from our existing study. We demonstrate our combined goodness measure in a number of applications, such as automatically selecting a good set of representative views, optimizing camera orbits to pass through good views and avoid bad views, and trackball controls that gently guide the viewer towards better views.
Adrian Secord, Jingwan Lu, Adam Finkelstein, Manish Singh 0001, Andrew Nealen
ACM Trans. Graph.3
2010 The Generalized PatchMatch Correspondence Algorithm
Connelly Barnes, Eli Shechtman, Dan B. Goldman, Adam Finkelstein
ECCV (3)4
2010 Interactive painterly stylization of images, videos and 3D animations
abstract
We introduce a real-time system that converts images, video, or 3D animation sequences to artistic renderings in various painterly styles. The algorithm, which is entirely executed on the GPU, can efficiently process 512 resolution frames containing 60,000 individual strokes at over 30 fps. In order to exploit the parallel nature of GPUs, our algorithm determines the placement of strokes entirely from local pixel neighborhood information. The strokes are rendered as point sprites with textures. Temporal coherence is achieved by treating the brush strokes as particles and moving them based on optical flow. Our system renders high quality results while allowing the user interactive control over many stylistic parameters such as stroke size, texture and density.
Jingwan Lu, Pedro V. Sander, Adam Finkelstein
SI3D3
2010 Sketcha: a captcha based on line drawings of 3D models
abstract
This paper introduces a captcha based on upright orientation of line drawings rendered from 3D models. The models are selected from a large database, and images are rendered from random viewpoints, affording many different drawings from a single 3D model. The captcha presents the user with a set of images, and the user must choose an upright orientation for each image. This task generally requires understanding of the semantic content of the image, which is believed to be difficult for automatic algorithms. We describe a process called covert filtering whereby the image database can be continually refreshed with drawings that are known to have a high success rate for humans, by inserting randomly into the captcha new images to be evaluated. Our analysis shows that covert filtering can ensure that captchas are likely to be solvable by humans while deterring attackers who wish to learn a portion of the database. We performed several user studies that evaluate how effectively people can solve the captcha. Comparing these results to an attack based on machine learning, we find that humans possess a substantial performance advantage over computers.
Steven A. Ross, J. Alex Halderman, Adam Finkelstein
WWW3
2010 Video tapestries with continuous temporal zoom
abstract
We present a novel approach for summarizing video in the form of a multiscale image that is continuous in both the spatial domain and across the scale dimension: There are no hard borders between discrete moments in time, and a user can zoom smoothly into the image to reveal additional temporal details. We call these artifacts tapestries because their continuous nature is akin to medieval tapestries and other narrative depictions predating the advent of motion pictures. We propose a set of criteria for such a summarization, and a series of optimizations motivated by these criteria. These can be performed as an entirely offline computation to produce high quality renderings, or by adjusting some optimization parameters the later stages can be solved in real time, enabling an interactive interface for video navigation. Our video tapestries combine the best aspects of two common visualizations, providing the visual clarity of DVD chapter menus with the information density and multiple scales of a video editing timeline representation. In addition, they provide continuous transitions between zoom levels. In a user study, participants preferred both the aesthetics and efficiency of tapestries over other interfaces for visual browsing.
Connelly Barnes, Dan B. Goldman, Eli Shechtman, Adam Finkelstein
ACM Trans. Graph.4
2010 Two Fast Methods for High-Quality Line Visibility
abstract
Lines drawn over or in place of shaded 3D models can often provide greater comprehensibility and stylistic freedom than shading alone. A substantial challenge for making stylized line drawings from 3D models is the visibility computation. Current algorithms for computing line visibility in models of moderate complexity are either too slow for interactive rendering, or too brittle for coherent animation. We introduce two methods that exploit graphics hardware to provide fast and robust line visibility. First, we present a simple shader that performs a visibility test for high-quality, simple lines drawn with the conventional implementation. Next, we offer a full optimized pipeline that supports line visibility and a broad range of stylization options.
Forrester Cole, Adam Finkelstein
IEEE Trans. Vis. Comput. Graph.2
2009 Fast high-quality line visibility
abstract
Lines drawn over or in place of shaded 3D models can often provide greater comprehensibility and stylistic freedom that shading alone. A substantial challenge for making stylized line drawings from 3D models is the visibility computation. Current algorithms for computing line visibility in models of moderate complexity are either too slow for interactive rendering, or too brittle for coherent animation. We present a method that exploits graphics hardware to provide fast and robust line visibility. Rendering speed for our system is usually within a factor of two of an optimized rendering pipeline using conventional lines, and our system provides much higher visual quality and flexibility for stylization.
Forrester Cole, Adam Finkelstein
SI3D2
2009 Fingerprinting Blank Paper Using Commodity Scanners
abstract
We develop a novel technique for authenticating physical documents by using random, naturally occurring imperfections in paper texture. To this end, we devised a new method for measuring the three-dimensional surface of a paper without modifying the document in any way, using only a commodity scanner. From this physical feature, we generate a concise fingerprint that uniquely identifies the document. Our method is secure against counterfeiting, robust to harsh handling, and applicable even before any content is printed on a page. It has a wide range of applications, including detecting forged currency and tickets, authenticating passports, and halting counterfeit goods. On a more sinister note, document identification could be used to de-anonymize printed surveys and to compromise the secrecy of paper ballots.
William Clarkson, Tim Weyrich, Adam Finkelstein, Nadia Heninger, J. Alex Halderman, Edward W. Felten
SP3
2009 PatchMatch: a randomized correspondence algorithm for structural image editing
abstract
This paper presents interactive image editing tools using a new randomized algorithm for quickly finding approximate nearest-neighbor matches between image patches. Previous research in graphics and vision has leveraged such nearest-neighbor searches to provide a variety of high-level digital image editing tools. However, the cost of computing a field of such matches for an entire image has eluded previous efforts to provide interactive performance. Our algorithm offers substantial performance improvements over the previous state of the art (20-100x), enabling its use in interactive editing tools. The key insights driving the algorithm are that some good patch matches can be found via random sampling, and that natural coherence in the imagery allows us to propagate such matches quickly to surrounding areas. We offer theoretical analysis of the convergence properties of the algorithm, as well as empirical and practical evidence for its high quality and performance. This one simple algorithm forms the basis for a variety of tools -- image retargeting, completion and reshuffling -- that can be used together in the context of a high-level image editing application. Finally, we propose additional intuitive constraints on the synthesis process that offer the user a level of control unavailable in previous methods.
Connelly Barnes, Eli Shechtman, Adam Finkelstein, Dan B. Goldman
ACM Trans. Graph.3
2009 How well do line drawings depict shape?
abstract
This paper investigates the ability of sparse line drawings to depict 3D shape. We perform a study in which people are shown an image of one of twelve 3D objects depicted with one of six styles and asked to orient a gauge to coincide with the surface normal at many positions on the object's surface. The normal estimates are compared with each other and with ground truth data provided by a registered 3D surface model to analyze accuracy and precision. The paper describes the design decisions made in collecting a large data set (275,000 gauge measurements) and provides analysis to answer questions about how well people interpret shapes from drawings. Our findings suggest that people interpret certain shapes almost as well from a line drawing as from a shaded image, that current computer graphics line drawing techniques can effectively depict shape and even match the effectiveness of artist's drawings, and that errors in depiction are often localized and can be traced to particular properties of the lines used. The data collected for this study will become a publicly available resource for further studies of this type.
Forrester Cole, Kevin Sanik, Douglas DeCarlo, Adam Finkelstein, Thomas A. Funkhouser, Szymon Rusinkiewicz, Manish Singh 0001
ACM Trans. Graph.4
2008 Video puppetry: a performative interface for cutout animation
abstract
We present a video-based interface that allows users of all skill levels to quickly create cutout-style animations by performing the character motions. The puppeteer first creates a cast of physical puppets using paper, markers and scissors. He then physically moves these puppets to tell a story. Using an inexpensive overhead camera our system tracks the motions of the puppets and renders them on a new background while removing the puppeteer's hands. Our system runs in real-time (at 30 fps) so that the puppeteer and the audience can immediately see the animation that is created. Our system also supports a variety of constraints and effects including articulated characters, multi-track animation, scene changes, camera controls, 2 1/2-D environments, shadows, and animation cycles. Users have evaluated our system both quantitatively and qualitatively: In tests of low-level dexterity, our system has similar accuracy to a mouse interface. For simple story telling, users prefer our system over either a mouse interface or traditional puppetry. We demonstrate that even first-time users, including an eleven-year-old, can use our system to quickly turn an original story idea into an animation.
Connelly Barnes, David E. Jacobs, Jason Sanders, Dan B. Goldman, Szymon Rusinkiewicz, Adam Finkelstein, Maneesh Agrawala
ACM Trans. Graph.6
2008 Adaptive cutaways for comprehensible rendering of polygonal scenes
abstract
In 3D renderings of complex scenes, objects of interest may be occluded by those of secondary importance. Cutaway renderings address this problem by omitting portions of secondary objects so as to expose the objects of interest. This paper introduces a method for generating cutaway renderings of polygonal scenes at interactive frame rates, using illustrative and non-photorealistic rendering cues to expose objects of interest in the context of surrounding objects. We describe a method for creating a view-dependent cutaway shape along with modifications to the polygonal rendering pipeline to create cutaway renderings. Applications for this technique include architectural modeling, path planning, and computer games.
Michael Burns, Adam Finkelstein
ACM Trans. Graph.2
2008 Where do people draw lines?
abstract
This paper presents the results of a study in which artists made line drawings intended to convey specific 3D shapes. The study was designed so that drawings could be registered with rendered images of 3D models, supporting an analysis of how well the locations of the artists' lines correlate with other artists', with current computer graphics line definitions, and with the underlying differential properties of the 3D surface. Lines drawn by artists in this study largely overlapped one another (75% are within 1mm of another line), particularly along the occluding contours of the object. Most lines that do not overlap contours overlap large gradients of the image intensity, and correlate strongly with predictions made by recent line drawing algorithms in computer graphics. 14% were not well described by any of the local properties considered in this study. The result of our work is a publicly available data set of aligned drawings, an analysis of where lines appear in that data set based on local properties of 3D models, and algorithms to predict where artists will draw lines for new scenes.
Forrester Cole, Aleksey Golovinskiy, Alex Limpaecher, Heather Stoddart Barros, Adam Finkelstein, Thomas A. Funkhouser, Szymon Rusinkiewicz
ACM Trans. Graph.5
2007 Lighting with paint
abstract
Lighting is a fundamental aspect of computer cinematography that involves the placement and configuration of lights to establish mood and enhance storytelling. This process is labor intensive as artists repeatedly adjust the parameters of a large set of complex lights to achieve a desired effect. Typical lighting controls affect the final image indirectly, requiring a large number of trials to obtain a suitable result. We present an interactive system wherein an artist paints desired lighting effects directly into the scene, and the computer solves for parameters that achieve the desired look. The artist can paint color, light shape, shadows, highlights, and reflections using a suite of tools designed for painting light. Our system matches these effects using a nonlinear optimizer made robust by a combination of initial estimates, system design, and user-guided optimization. In contrast, previous work on painting light has not permitted the lights to move, allowing for linear optimization but preventing its use in computer cinematography. To demonstrate our approach we lit several scenes, mainly using a direct illumination renderer designed for computer animation, but also including two other rendering styles. We show that painting interfaces can quickly produce high quality lighting setups, easing the lighting artist's workflow.
Fabio Pellacini, Frank Battaglia, R. Keith Morley, Adam Finkelstein
ACM Trans. Graph.4
2007 Digital bas-relief from 3D scenes
abstract
We present a system for semi-automatic creation of bas-relief sculpture. As an artistic medium, relief spans the continuum between 2D drawing or painting and full 3D sculpture. Bas-relief (or low relief) presents the unique challenge of squeezing shapes into a nearly-flat surface while maintaining as much as possible the perception of the full 3D scene. Our solution to this problem adapts methods from the tone-mapping literature, which addresses the similar problem of squeezing a high dynamic range image into the (low) dynamic range available on typical display devices. However, the bas-relief medium imposes its own unique set of requirements, such as maintaining small, fixed-size depth discontinuities. Given a 3D model, camera, and a few parameters describing the relative attenuation of different frequencies in the shape, our system creates a relief that gives the illusion of the 3D shape from a given vantage point while conforming to a greatly compressed height.
Tim Weyrich, Jia Deng 0001, Connelly Barnes, Szymon Rusinkiewicz, Adam Finkelstein
ACM Trans. Graph.5
2006 Directing Gaze in 3D Models with Stylized Focus
Forrester Cole, Douglas DeCarlo, Adam Finkelstein, Kenrick Kin, R. Keith Morley, Anthony Santella
Rendering Techniques3
2005 Line drawings from volume data
abstract
Renderings of volumetric data have become an important data analysis tool for applications ranging from medicine to scientific simulation. We propose a volumetric drawing system that directly extracts sparse linear features, such as silhouettes and suggestive contours, using a temporally coherent seed-and-traverse framework. In contrast to previous methods based on isosurfaces or nonrefractive transparency, producing these drawings requires examining an asymptotically smaller subset of the data, leading to efficiency on large data sets. In addition, the resulting imagery is often more comprehensible than standard rendering styles, since it focuses attention on important features in the data. We test our algorithms on datasets up to 512 3 , demonstrating interactive extraction and rendering of line drawings in a variety of drawing styles.
Michael Burns, Janek Klawe, Szymon Rusinkiewicz, Adam Finkelstein, Douglas DeCarlo
ACM Trans. Graph.4
2003 Suggestive contours for conveying shape
abstract
In this paper, we describe a non-photorealistic rendering system that conveys shape using lines. We go beyond contours and creases by developing a new type of line to draw: the suggestive contour . Suggestive contours are lines drawn on clearly visible parts of the surface, where a true contour would first appear with a minimal change in viewpoint. We provide two methods for calculating suggestive contours, including an algorithm that finds the zero crossings of the radial curvature. We show that suggestive contours can be drawn consistently with true contours, because they anticipate and extend them. We present a variety of results, arguing that these images convey shape more effectively than contour alone.
Douglas DeCarlo, Adam Finkelstein, Szymon Rusinkiewicz, Anthony Santella
ACM Trans. Graph.2
2003 Coherent stylized silhouettes
abstract
We describe a way to render stylized silhouettes of animated 3D models with temporal coherence. Coherence is one of the central challenges for non-photorealistic rendering. It is especially difficult for silhouettes, because they may not have obvious correspondences between frames. We demonstrate various coherence effects for stylized silhouettes with a robust working system. Our method runs in real-time for models of moderate complexity, making it suitable for both interactive applications and offline animation.
Robert D. Kalnins, Phillip L. Davidson, Lee Markosian, Adam Finkelstein
ACM Trans. Graph.4
2002 A Reflective Symmetry Descriptor
Michael M. Kazhdan, Bernard Chazelle, David P. Dobkin, Adam Finkelstein, Thomas A. Funkhouser
ECCV (2)4
2002 Improving progressive view-dependent isosurface propagation
Zhiyan Liu, Adam Finkelstein, Kai Li 0001
Comput. Graph.2
2002 WYSIWYG NPR: drawing strokes directly on 3D models
abstract
We present a system that lets a designer directly annotate a 3D model with strokes, imparting a personal aesthetic to the non-photorealistic rendering of the object. The artist chooses a "brush" style, then draws strokes over the model from one or more viewpoints. When the system renders the scene from any new viewpoint, it adapts the number and placement of the strokes appropriately to maintain the original look.
Robert D. Kalnins, Lee Markosian, Barbara J. Meier, Michael A. Kowalski, Joseph C. Lee, Phillip L. Davidson, Matthew Webb, John F. Hughes, Adam Finkelstein
ACM Trans. Graph.9
2002 A framework for geometric warps and deformations
abstract
We present a framework for geometric warps and deformations. The framework provides a conceptual and mathematical foundation for analyzing known warps and for developing new warps, and serves as a common base for many warps and deformations. Our framework is composed of two components: a generic modular algorithm for warps and deformations; and a concise, geometrically meaningful formula that describes how warps are evaluated. Together, these two elements comprise a complete framework useful for analyzing, evaluating, designing, and implementing deformation algorithms. While the framework is independent of user-interfaces and geometric model representations and is formally capable of describing any warping algorithm, its design is geared toward the most prevalent class of user-controlled deformations: those computed using geometric operations. To demonstrate the expressive power of the framework, we cast several well-known warps in terms of the framework. To illustrate the framework's usefulness for analyzing and modifying existing warps, we present variations of these warps that provide additional functionality or improved behavior. To show the utility of the framework for developing new warps, we design a novel 3-D warping algorithm: a mesh warp ---useful as a modeling and animation tool---that allows users to deform a detailed surface by manipulating a low-resolution mesh of similar shape. Finally, to demonstrate the mathematical utility of the framework, we use the framework to develop guarantees of several mathematical properties such as commutativity and continuity for large classes of deformations.
Tim Milliron, Robert J. Jensen, Ronen Barzel, Adam Finkelstein
ACM Trans. Graph.4
2001 Real-time fur over arbitrary surfaces
abstract
We introduce a method for real-time rendering of fur on surfaces of arbitrary topology. As a pre-process, we simulate virtual hair with a particle system, and sample it into a volume texture. Next, we parameterize the texture over a surface of arbitrary topology using "lapped textures" --- an approach for applying a sample texture to a surface by repeatedly pasting patches of the texture until the surface is covered. The use of lapped textures permits specifying a global direction field for the fur over the surface. At runtime, the patches of volume textures are rendered as a series of concentric shells of semi-transparent medium. To improve the visual quality of the fur near silhouettes, we place "fins" normal to the surface and render these using conventional 2D texture maps sampled from the volume texture in the direction of hair growth. The method generates convincing imagery of fur at interactive rates for models of moderate complexity. Furthermore, the scheme allows real-time modification of viewing and lighting conditions, as well as local control over hair color, length, and direction.
Jed Lengyel, Emil Praun, Adam Finkelstein, Hugues Hoppe
SI3D3
2001 Real-time hatching
abstract
Drawing surfaces using hatching strokes simultaneously conveys material, tone, and form. We present a real-time system for non-photorealistic rendering of hatching strokes over arbitrary surfaces. During an automatic preprocess, we construct a sequence of mipmapped hatch images corresponding to different tones, collectively called a tonal art map. Strokes within the hatch images are scaled to attain appropriate stroke size and density at all resolutions, and are organized to maintain coherence across scales and tones. At runtime, hardware multitexturing blends the hatch images over the rendered faces to locally vary tone while maintaining both spatial and temporal coherence. To render strokes over arbitrary surfaces, we build a lapped texture parametrization where the overlapping patches align to a curvature-based direction field. We demonstrate hatching strokes over complex surfaces in a variety of styles.
Emil Praun, Hugues Hoppe, Matthew Webb, Adam Finkelstein
SIGGRAPH4
2001 Data distribution strategies for high-resolution displays
Yuqun Chen, Adam Finkelstein, Thomas A. Funkhouser, Kai Li 0001, Zhiyan Liu, Rudrajit Samanta, Grant Wallace
Comput. Graph.3
2000 Non-photorealistic virtual environments
abstract
We describe a system for non-photorealistic rendering (NPR) of virtual environments. In real time, it synthesizes imagery of architectural interiors using stroke-based textures. We address the four main challenges of such a system — interactivity, visual detail, controlled stroke size, and frame-to-frame coherence — through image based rendering (IBR) methods. In a preprocessing stage, we capture photos of a real or synthetic environment, map the photos to a coarse model of the environment, and run a series of NPR filters to generate textures. At runtime, the system re-renders the NPR textures over the geometry of the coarse model, and it adds dark lines that emphasize creases and silhouettes. We provide a method for constructing non-photorealistic textures from photographs that largely avoids seams in the resulting imagery. We also offer a new construction, art-maps, to control stroke size across the images. Finally, we show a working system that provides an immersive experience rendered in a variety of NPR styles.
Allison W. Klein, Wilmot Li, Michael M. Kazhdan, Wagner Toledo Corrêa, Adam Finkelstein, Thomas A. Funkhouser
SIGGRAPH5
2000 Shadows for cel animation
abstract
We present a semi-automatic method for creating shadow mattes in cel animation. In conventional cel animation, shadows are drawn by hand, in order to provide visual cues about the spatial relationships and forms of characters in the scene. Our system creates shadow mattes based on hand-drawn characters, given high-level guidance from the user about depths of various objects. The method employs a scheme for "inflating" a 3D figure based on hand-drawn art. It provides simple tools for adjusting object depths, coupled with an intuitive interface by which the user specifies object shapes and relative positions in a scene. Our system obviates the tedium of drawing shadow mattes by hand, and provides control over complex shadows falling over interesting shapes. Keywords: Shadows, cel animation, inflation, sketching, NPR. URL: http://www.cs.princeton.edu/gfx/proj/cel shadows 1 Introduction Shadows provide important visual cues for depth, shape, contact, movement, and lighting in our perce...
Lena Petrovic, Brian Fujito, Lance Williams, Adam Finkelstein
SIGGRAPH4
2000 Lapped textures
abstract
We present for creating texture over an surface mesh using an example 2D texture. The approach is to identify interesting regions (texture patches) in the 2D example, and to repeatedly paste them onto the surface until it is completely covered. We call such a collection of overlapping patches a lapped texture. It is rendered using compositing operations, either into a traditional global texture map during a preprocess, or directly with the surface at runtime. The runtime compositing approach avoids resampling artifacts and drastically reduces texture memory requirements.
Emil Praun, Adam Finkelstein, Hugues Hoppe
SIGGRAPH2
2000 Automatic alignment of high-resolution multi-projector display using an un-calibrated camera
abstract
A scalable, high-resolution display may be constructed by tiling many projected images over a single display surface. One fundamental challenge for such a display is to avoid visible seams due to misalignment among the projectors. Traditional methods for avoiding seams involve sophisticated mechanical devices and expensive CRT projectors, coupled with extensive human effort for fine-tuning the projectors. The paper describes an automatic alignment method that relies on an inexpensive, uncalibrated camera to measure the relative mismatches between neighboring projectors, and then correct the projected imagery to avoid seams without significant human effort.
Yuqun Chen, Douglas W. Clark, Adam Finkelstein, Timothy C. Housel, Kai Li 0001
IEEE Visualization3
1999 Robust Mesh Watermarking
abstract
Watermarking provides a mechanism for copyright protection of digital media by embedding infor-mation identifying the owner in the data. The bulk of the research on digital watermarks has focused on media such as images, video, audio, and text. Robust watermarks must be able to survive a va-riety of “attacks”, including resizing, cropping, and filtering. For resilience to such attacks, recent watermarking schemes employ a “spread-spectrum ” approach — they transform the document to the frequency domain (e.g. using DCT) and perturb the coefficients of the perceptually most significant basis functions. In this paper we extend this spread-spectrum approach for the robust watermarking of arbitrary triangle meshes. Generalization of the spread spectrum techniques to surfaces presents two major challenges. First, arbitrary surfaces lack a natural parametrization for frequency-based decomposition. Our solution is to construct a set of scalar basis function over the mesh vertices using a multiresolution analysis of the mesh. The watermark is embedded in the mesh by perturbing vertices along the direction of the surface normal, weighted by the basis functions. The second challenge is that attacks such as simplification may modify the connectivity of the mesh. We use an optimization technique to resam-ple an attacked mesh using the original mesh connectivity. Results demonstrate that our watermarks are resistant to common mesh processing operations such as translation, rotation, scaling, cropping, smoothing, simplification, and resampling, as well as malicious attacks such as the insertion of noise, modification of low-order bits, or even insertion of other watermarks.
Emil Praun, Hugues Hoppe, Adam Finkelstein
SIGGRAPH3
1999 Wavelet-Based Video Indexing and Querying
Xiaodong Wen, Theodore D. Huffmire, Helen H. Hu, Adam Finkelstein
Multim. Syst.4
1998 Texture Mapping for Cell Animation
abstract
We present a method for applying complex textures to hand-drawn characters in cel animation. The method correlates features in a simple, textured, 3-D model with features on a hand-drawn figure, and then distorts the model to conform to the hand-drawn artwork. The process uses two new algorithms: a silhouette detection scheme and a depth-preserving warp. The silhouette detection algorithm is simple and efficient, and it produces continuous, smooth, visible contours on a 3-D model. The warp distorts the model in only two dimensions to match the artwork from a given camera perspective, yet preserves 3-D effects such as self-occlusion and foreshortening. The entire process allows animators to combine complex textures with hand-drawn artwork, leveraging the strengths of 3-D computer graphics while retaining the expressiveness of traditional handdrawn cel animation.
Wagner Toledo Corrêa, Robert J. Jensen, Craig E. Thayer, Adam Finkelstein
SIGGRAPH4
1997 Multiperspective panoramas for cel animation
abstract
We describe a new approach for simulating apparent camera motion through a 3D environment. The approach is motivated by a traditional technique used in 2D cel animation, in which a single background image, which we call a multiperspective panorama,is used to incorporate multiple views of a 3D environment as seen from along a given camera path. When viewed through a small moving window, the panorama produces the illusion of 3D motion. In this paper, we explore how such panoramas can be designed by computer, and we examine their application to cel animation in particular. Multiperspective panoramas should also be useful for any application in which predefined camera moves are applied to 3D scenes, including virtual reality fly-throughs, computer games, and architectural walk-throughs. CR Categories: I.3.3 [Computer Graphics]: Picture/Image Generation. Additional Keywords: CGI production, compositing, illustration, imagebased rendering, mosaics, multiplaning, non-photorealistic renderi...
Daniel N. Wood, Adam Finkelstein, John F. Hughes, Craig E. Thayer, David Salesin
SIGGRAPH2
1996 Multiresolution Video
abstract
We present a new representation for time-varying image data that allows for varying-and arbitrarily high-spatial and temporal resolutions in different parts of a video sequence.The representation, called multiresolution video, is based on a sparse, hierarchical encoding of the video data.We describe a number of operations for creating, viewing, and editing multiresolution sequences.These operations support a variety of applications: multiresolution playback, including motion-blurred "fast-forward" and "reverse"; constantspeed display; enhanced video scrubbing; and "video clip-art" editing and compositing.The multiresolution representation requires little storage overhead, and the algorithms using the representation are both simple and efficient.
Adam Finkelstein, Charles E. Jacobs, David Salesin
SIGGRAPH1
1995 Fast multiresolution image querying
abstract
We present a method for searching in an image database using a query image that is similar to the intended target.The query image may be a hand-drawn sketch or a (potentially low-quality) scan of the image to be retrieved.Our searching algorithm makes use of multiresolution wavelet decompositions of the query and database images.The coefficients of these decompositions are distilled into small "signatures" for each image.We introduce an "image querying metric" that operates on these signatures.This metric essentially compares how many significant wavelet coefficients the query has in common with potential targets.The metric includes parameters that can be tuned, using a statistical analysis, to accommodate the kinds of image distortions found in different types of image queries.The resulting algorithm is simple, requires very little storage overhead for the database of signatures, and is fast enough to be performed on a database of 20,000 images at interactive rates (on standard desktop machines) as a query is sketched.Our experiments with hundreds of queries in databases of 1000 and 20,000 images show dramatic improvement, in both speed and success rate, over using a conventional L 1 , L 2 , or color histogram norm.
Charles E. Jacobs, Adam Finkelstein, David Salesin
SIGGRAPH2
1994 Multiresolution curves
abstract
We describe a multiresolution curve representation, based on wavelets, that conveniently supports a variety of operations: smoothing a curve; editing the overall form of a curve while preserving its details; and approximating a curve within any given error tolerance for scan conversion. We present methods to support continuous levels of smoothing as well as direct manipulation of an arbitrary portion of the curve; the control points, as well as the discrete nature of the underlying hierarchical representation, can be hidden from the user. The multiresolution representation requires no extra storage beyond that of the original control points, and the algorithms using the representation are both simple and fast.
Adam Finkelstein, David Salesin
SIGGRAPH1
1993 Real-Time Self-Explanatory Simulation
Franz G. Amador, Adam Finkelstein, Daniel S. Weld
AAAI2
1993 Electronic "How Things Work" Articles: Two Early Prototypes
abstract
The electronic encyclopedia exploratorium (E/sup 3/) is a vision of a future computer system-an electronic book describing how thing work. Typical articles in E/sup 3/ will describe such mechanisms as compression refrigerators, engines, telescopes, and mechanical linkages. Each article will provide simulations, three-dimensional animated graphics that the user can manipulate, laboratory areas that allow a user to modify the device or experiment with related artifacts, and a facility for asking questions and receiving customized, computer-generated English-language explanations. Some of the foundational technology is discussed, focusing on topics in artificial intelligence, graphics, and user interfaces. The initial prototype system and the technical lessons learned from it, as well as the second prototype currently under construction, are described.>
Franz G. Amador, Deborah Berman, Alan Borning, Tony DeRose, Adam Finkelstein, Dorothy Neville, David Notkin, David Salesin, Michael Salisbury, Joe Sherman, Daniel S. Weld, Georges Winkenbach
IEEE Trans. Knowl. Data Eng.5