EDBT 2026 Demo / reviewers in the wild / expert
Steven M. Seitz
dblp:s/StevenMSeitz
· DBLP profile ↗
135ranked-venue papers
16as first author
18since 2021 · last 2025
0009-0000-4214-4078ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 103 · 11 first-author · 13 since 2021Artificial intelligence and machine learning · 92 · 15 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 14 · 1 first-author · 4 since 2021Computer networks · 3 · 2 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Generative Inbetweening: Adapting Image-to-Video Models for Keyframe InterpolationabstractWe present a method for generating video sequences with coherent motion between a pair of input keyframes. We adapt a pretrained large-scale image-to-video diffusion model (originally trained to generate videos moving forward in time from a single input image) for keyframe interpolation, i.e., to produce a video between two input frames. We accomplish this adaptation through a lightweight fine-tuning technique that produces a version of the model that instead predicts videos moving backwards in time from a single input image. This model (along with the original forward-moving model) is subsequently used in a dual-directional diffusion sampling process that combines the overlapping model estimates starting from each of the two keyframes. Our experiments shows that our method outperforms both existing diffusion-based methods and traditional frame interpolation techniques. Brian Curless, Ira Kemelmacher-Shlizerman, Aleksander Holynski, Steven M. Seitz |
ICLR | 6 |
| 2025 | Linearly Constrained Diffusion Implicit ModelsabstractWe introduce Linearly Constrained Diffusion Implicit Models (CDIM), a fast and accurate approach to solving noisy linear inverse problems using diffusion models. Traditional diffusion-based inverse methods rely on numerous projection steps to enforce measurement consistency in addition to unconditional denoising steps. CDIM achieves a 10–50× reduction in projection steps by dynamically adjusting the number and size of projection steps to align a residual measurement energy with its theoretical distribution under the forward diffusion process. This adaptive alignment preserves measurement consistency while substantially accelerating constrained inference.
For noise-free linear inverse problems, CDIM exactly satisfies the measurement constraints with few projection steps, even when existing methods fail. We demonstrate CDIM’s effectiveness across a range of applications, including super-resolution, denoising, inpainting, deblurring, and 3D point cloud reprojection. Vivek Jayaram, Ira Kemelmacher-Shlizerman, Steven M. Seitz, John Thickstun |
NeurIPS | 3 |
| 2025 | UltraZoom: Generating Gigapixel Images from Regular PhotosabstractWe present UltraZoom, a system for generating gigapixel-resolution images of objects from casually captured inputs, such as handheld phone photos. Given a full-shot image (global, low-detail) and one or more close-ups (local, high-detail), UltraZoom upscales the full image to match the fine detail and scale of the close-up examples. To achieve this, we construct a per-instance paired dataset from the close-ups and adapt a pretrained generative model to learn object-specific low-to-high resolution mappings. At inference, we apply the model in a sliding window fashion over the full image. Constructing these pairs is non-trivial: it requires registering the close-ups within the full image for scale estimation and degradation alignment. We introduce a simple, robust method for achieving registration on arbitrary materials in casual, in-the-wild captures. Together, these components form a system that enables seamless pan and zoom across the entire object, producing consistent, photorealistic gigapixel imagery from minimal input. For full-resolution results and code, visit our project page at ultra-zoom.github.io . Vivek Jayaram, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 5 |
| 2024 | Total Selfie: Generating Full-Body SelfiesabstractWe present a method to generate full-body selfies from photographs originally taken at arms length. Because self-captured photos are typically taken close up, they have lim-ited field of view and exaggerated perspective that distorts facial shapes. We instead seek to generate the photo some one else would take of you from a few feet away. Our approach takes as input four selfies of your face and body, a background image, and generates a full-body selfie in a de-sired target pose. We introduce a novel diffusion-based approach to combine all of this information into high-quality, well-composed photos of you with the desired pose and background. Bowei Chen 0003, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
CVPR | 4 |
| 2024 | Generative Powers of TenabstractWe present a method that uses a text-to-image model to generate consistent content across multiple image scales, enabling extreme semantic zooms into a scene, e.g. ranging from a wide-angle landscape view of a forest to a macro shot of an insect sitting on one of the tree branches. We achieve this through a joint multi-scale diffusion sampling approach that encourages consistency across different scales while preserving the integrity of each individual sampling process. Since each generated scale is guided by a different text prompt, our method enables deeper levels of zoom than traditional super-resolution methods that may struggle to create new contextual structure at vastly different scales. We compare our method qualitatively with alter-native techniques in image super-resolution and outpainting, and show that our method is most effective at generating consistent multi-scale content. Janne Kontkanen, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman, Ben Mildenhall, Pratul P. Srinivasan, Dor Verbin, Aleksander Holynski |
CVPR | 4 |
| 2024 | Inverse Painting: Reconstructing The Painting Process
Bowei Chen 0003, Yifan Wang 0013, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 5 |
| 2023 | Animating Street ViewabstractWe present a system that automatically brings street view imagery to life by populating it with naturally behaving, animated pedestrians and vehicles. Our approach is to remove existing people and vehicles from the input image, insert moving objects with proper scale, angle, motion and appearance, plan paths and traffic behavior, as well as render the scene with plausible occlusion and shadowing effects. The system achieves these by reconstructing the still image street scene, simulating crowd behavior, and rendering with consistent lighting, visibility, occlusions, and shadows. We demonstrate results on a diverse range of street scenes including regular still images and panoramas. Mengyi Shan, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
SIGGRAPH Asia | 4 |
| 2023 | HRTF Estimation in the WildabstractHead Related Transfer Functions (HRTFs) play a crucial role in creating immersive spatial audio experiences. However, HRTFs differ significantly from person to person, and traditional methods for estimating personalized HRTFs are expensive, time-consuming, and require specialized equipment. We imagine a world where your personalized HRTF can be determined by capturing data through earbuds in everyday environments. In this paper, we propose a novel approach for deriving personalized HRTFs that only relies on in-the-wild binaural recordings and head tracking data. By analyzing how sounds change as the user rotates their head through different environments with different noise sources, we can accurately estimate their personalized HRTF. Our results show that our predicted HRTFs closely match ground-truth HRTFs measured in an anechoic chamber. Furthermore, listening studies demonstrate that our personalized HRTFs significantly improve sound localization and reduce front-back confusion in virtual environments. Our approach offers an efficient and accessible method for deriving personalized HRTFs and has the potential to greatly improve spatial audio experiences. Vivek Jayaram, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
UIST | 3 |
| 2022 | ClearBuds: wireless binaural earbuds for learning-based speech enhancementabstractWe present ClearBuds, the first hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our neural network has a novel cascaded architecture that combines a time-domain conventional neural network with a spectrogram-based frequency masking neural network to reduce the artifacts in the audio output. Results show that our wireless earbuds achieve a synchronization error less than 64 μs and our network has a runtime of 21.4 ms on an accompanying mobile phone. In-the-wild evaluation with eight users in previously unseen indoor and outdoor multipath scenarios demonstrates that our neural network generalizes to learn both spatial and acoustic cues to perform noise suppression and background speech removal. In a user-study with 37 participants who spent over 15.4 hours rating 1041 audio samples collected in-the-wild, our system achieves improved mean opinion score and background noise suppression. Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher-Shlizerman, Shwetak N. Patel, Steven M. Seitz |
MobiSys | 7 |
| 2022 | ClearBuds - wireless binaural earbuds for learning-based speech enhancementabstractWe present ClearBuds, the first end-to-end hardware and software system that utilizes a neural network to enhance speech streamed from two wireless earbuds. Real-time speech enhancement for wireless earbuds requires high-quality sound separation and background cancellation, operating in real-time and on a mobile phone. Clear-Buds bridges state-of-the-art deep learning for blind audio source separation and in-ear mobile systems by making two key technical contributions: 1) a new wireless earbud design capable of operating as a synchronized, binaural microphone array, and 2) a lightweight dual-channel speech enhancement neural network that runs on a mobile device. Our demo will allow MobiSys attendees wear our earbuds, and experience noise suppression as they talk in a noisy environment. Companion video can be accessed using the link below: Ishan Chatterjee, Maruchi Kim, Vivek Jayaram, Shyamnath Gollakota, Ira Kemelmacher-Shlizerman, Shwetak N. Patel, Steven M. Seitz |
MobiSys | 7 |
| 2021 | Animating Pictures With Eulerian Motion FieldsabstractIn this paper, we demonstrate a fully automatic method for converting a still image into a realistic animated looping video. We target scenes with continuous fluid motion, such as flowing water and billowing smoke. Our method relies on the observation that this type of natural motion can be convincingly reproduced from a static Eulerian motion description, i.e. a single, temporally constant flow field that defines the immediate motion of a particle at a given 2D location. We use an image-to-image translation network to encode motion priors of natural scenes collected from on-line videos, so that for a new photo, we can synthesize a corresponding motion field. The image is then animated using the generated motion through a deep warping technique: pixels are encoded as deep features, those features are warped via Eulerian motion, and the resulting warped feature maps are decoded as images. In order to produce continuous, seamlessly looping video textures, we propose a novel video looping technique that flows features both for-ward and backward in time and then blends the results. We demonstrate the effectiveness and robustness of our method by applying it to a large collection of examples including beaches, waterfalls, and flowing rivers. Aleksander Holynski, Brian Curless, Steven M. Seitz, Richard Szeliski |
CVPR | 3 |
| 2021 | Real-Time High-Resolution Background MattingabstractWe introduce a real-time, high-resolution background replacement technique which operates at 30fps in 4K resolution, and 60fps for HD on a modern GPU. Our technique is based on background matting, where an additional frame of the background is captured and used in recovering the alpha matte and the foreground layer. The main challenge is to compute a high-quality alpha matte, preserving strand-level hair details, while processing high-resolution images in real-time. To achieve this goal, we employ two neural networks; a base network computes a low-resolution result which is refined by a second network operating at high-resolution on selective patches. We introduce two large-scale video and image matting datasets: VideoMatte240K and PhotoMatte13K/85. Our approach yields higher quality results compared to the previous state-of-the-art in background matting, while simultaneously yielding a dramatic boost in both speed and resolution. Shanchuan Lin, Andrey Ryabtsev, Roni Sengupta, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 5 |
| 2021 | Repopulating Street ScenesabstractWe present a framework for automatically reconfiguring images of street scenes by populating, depopulating, or repopulating them with objects such as pedestrians or vehicles. Applications of this method include anonymizing images to enhance privacy, generating data augmentations for perception tasks like autonomous driving, and composing scenes to achieve a certain ambiance, such as empty streets in the early morning. At a technical level, our work has three primary contributions: (1) a method for clearing images of objects, (2) a method for estimating sun direction from a single image, and (3) a way to compose objects in scenes that respects scene geometry and illumination. Each component is learned from data with minimal ground truth annotations, by making creative use of large-numbers of short image bursts of street scenes. We demonstrate convincing results on a range of street scenes and illustrate potential applications. Yifan Wang 0013, Andrew Liu 0001, Richard Tucker 0001, Jiajun Wu 0001, Brian Curless, Steven M. Seitz, Noah Snavely |
CVPR | 6 |
| 2021 | Nerfies: Deformable Neural Radiance FieldsabstractWe present the first method capable of photorealistically reconstructing deformable scenes using photos/videos captured casually from mobile phones. Our approach augments neural radiance fields (NeRF) by optimizing an additional continuous volumetric deformation field that warps each observed point into a canonical 5D NeRF. We observe that these NeRF-like deformation fields are prone to local minima, and propose a coarse-to-fine optimization method for coordinate-based models that allows for more robust optimization. By adapting principles from geometry processing and physical simulation to NeRF-like models, we propose an elastic regularization of the deformation field that further improves robustness. We show that our method can turn casually captured selfie photos/videos into deformable NeRF models that allow for photorealistic renderings of the subject from arbitrary viewpoints, which we dub "nerfies." We evaluate our method by collecting time-synchronized data using a rig with two mobile phones, yielding train/validation images of the same pose at different viewpoints. We show that our method faithfully reconstructs non-rigidly deforming scenes and reproduces unseen views with high fidelity. Keunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Steven M. Seitz, Ricardo Martin-Brualla |
ICCV | 6 |
| 2021 | A Light Stage on Every DeskabstractEvery time you sit in front of a TV or monitor, your face is actively illuminated by time-varying patterns of light. This paper proposes to use this time-varying illumination for synthetic relighting of your face with any new illumination condition. In doing so, we take inspiration from the light stage work of Debevec et al. [4], who first demonstrated the ability to relight people captured in a controlled lighting environment. Whereas existing light stages require expensive, room-scale spherical capture gantries and exist in only a few labs in the world, we demonstrate how to acquire useful data from a normal TV or desktop monitor. Instead of subjecting the user to uncomfortable rapidly flashing light patterns, we operate on images of the user watching a YouTube video or other standard content. We train a deep network on images plus monitor patterns of a given user and learn to predict images of that user under any target illumination (monitor pattern). Experimental evaluation shows that our method produces realistic relighting results. Video results are available at grail.cs.washington.edu/projects/Light_Stage_on_Every_Desk/. Roni Sengupta, Brian Curless, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ICCV | 4 |
| 2021 | Project starline: a high-fidelity telepresence systemabstractWe present a real-time bidirectional communication system that lets two people, separated by distance, experience a face-to-face conversation as if they were copresent. It is the first telepresence system that is demonstrably better than 2D videoconferencing, as measured using participant ratings (e.g., presence, attentiveness, reaction-gauging, engagement), meeting recall, and observed nonverbal behaviors (e.g., head nods, eyebrow movements). This milestone is reached by maximizing audiovisual fidelity and the sense of copresence in all design elements, including physical layout, lighting, face tracking, multi-view capture, microphone array, multi-stream compression, loudspeaker output, and lenticular display. Our system achieves key 3D audiovisual cues (stereopsis, motion parallax, and spatialized audio) and enables the full range of communication cues (eye contact, hand gestures, and body language), yet does not require special glasses or body-worn microphones/headphones. The system consists of a head-tracked autostereoscopic display, high-resolution 3D capture and rendering subsystems, and network transmission using compressed color and depth video streams. Other contributions include a novel image-based geometry fusion algorithm, free-space dereverberation, and talker localization. Jason Lawrence, Dan B. Goldman, Supreeth Achar, Gregory Major Blascovich, Joseph G. Desloge, Tommy Fortes, Eric M. Gomez, Sascha Häberling, Hugues Hoppe, Andy Huibers, Claude Knaus, Brian Kuschak, Ricardo Martin-Brualla, Harris Nover, Andrew Ian Russell, Steven M. Seitz, Kevin Tong |
ACM Trans. Graph. | 16 |
| 2021 | Time-travel rephotographyabstractMany historical people were only ever captured by old, faded, black and white photos, that are distorted due to the limitations of early cameras and the passage of time. This paper simulates traveling back in time with a modern camera to rephotograph famous subjects. Unlike conventional image restoration filters which apply independent operations like denoising, colorization, and superresolution, we leverage the StyleGAN2 framework to project old photos into the space of modern high-resolution photos, achieving all of these effects in a unified framework. A unique challenge with this approach is retaining the identity and pose of the subject in the original photo, while discarding the many artifacts frequently seen in low-quality antique photos. Our comparisons to current state-of-the-art restoration filters show significant improvements and compelling results for a variety of important historical people. Please go to time-travell-rephotography.github.io for many more results. Xuaner Cecilia Zhang, Paul Yoo, Ricardo Martin-Brualla, Jason Lawrence, Steven M. Seitz |
ACM Trans. Graph. | 6 |
| 2021 | HyperNeRF: a higher-dimensional representation for topologically varying neural radiance fieldsabstractNeural Radiance Fields (NeRF) are able to reconstruct scenes with unprecedented fidelity, and various recent works have extended NeRF to handle dynamic scenes. A common approach to reconstruct such non-rigid scenes is through the use of a learned deformation field mapping from coordinates in each input image into a canonical template coordinate space. However, these deformation-based approaches struggle to model changes in topology, as topological changes require a discontinuity in the deformation field, but these deformation fields are necessarily continuous. We address this limitation by lifting NeRFs into a higher dimensional space, and by representing the 5D radiance field corresponding to each individual input image as a slice through this "hyper-space". Our method is inspired by level set methods, which model the evolution of surfaces as slices through a higher dimensional surface. We evaluate our method on two tasks: (i) interpolating smoothly between "moments", i.e., configurations of the scene, seen in the input images while maintaining visual plausibility, and (ii) novel-view synthesis at fixed moments. We show that our method, which we dub HyperNeRF , outperforms existing methods on both tasks. Compared to Nerfies, HyperNeRF reduces average error rates by 4.1% for interpolation and 8.6% for novel-view synthesis, as measured by LPIPS. Additional videos, results, and visualizations are available at hypernerf.github.io. Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B. Goldman, Ricardo Martin-Brualla, Steven M. Seitz |
ACM Trans. Graph. | 8 |
| 2020 | KeystoneDepth: History in 3DabstractThis paper introduces KeystoneDepth, the largest and most diverse collection of rectified historical stereo image pairs to date, consisting of tens of thousands of stereographs of people, events, objects, and scenes recorded between 1864 and 1966. Leveraging the Keystone-Mast Collection of stereographs from the California Museum of Photography, we apply multiple processing steps to produce clean stereo image pairs, complete with calibration data, rectification transforms, and disparity maps. We introduce a novel stereo rectification technique based on the unique properties of antique stereo cameras. To better visualize the results on 2D displays, we also introduce a self-supervised deep view synthesis technique trained on historical imagery. Our dataset is available at http://keystonedepth.cs.washington.edu/. Yanmeng Kong, Jason Lawrence, Ricardo Martin-Brualla, Steven M. Seitz |
3DV | 5 |
| 2020 | Scene Recomposition by Learning-Based ICPabstractBy moving a depth sensor around a room, we compute a 3D CAD model of the environment, capturing the room shape and contents such as chairs, desks, sofas, and tables. Rather than reconstructing geometry, we match, place, and align each object in the scene to thousands of CAD models of objects. In addition to the fully automatic system, the key technical contribution is a novel approach for aligning CAD models to 3D scans, based on deep reinforcement learning. This approach, which we call Learning-based ICP, outperforms prior ICP methods in the literature, by learning the best points to match and conditioning on object viewpoint. LICP learns to align using only synthetic data and does not require ground truth annotation of object pose or keypoint pair matching in real scene scans. While LICP is trained on synthetic data and without 3D real scene annotations, it outperforms both learned local deep feature matching and geometric based alignment methods in real scenes. The proposed method is evaluated on real scenes datasets of SceneNN and ScanNet as well as synthetic scenes of SUNCG. High quality results are demonstrated on a range of real world scenes, with robustness to clutter, viewpoint, and occlusion. Hamid Izadinia, Steven M. Seitz |
CVPR | 2 |
| 2020 | Seeing the World in a Bag of ChipsabstractWe address the dual problems of novel view synthesis and environment reconstruction from hand-held RGBD sensors. Our contributions include 1) modeling highly specular objects, 2) modeling inter-reflections and Fresnel effects, and 3) enabling surface light field reconstruction with the same input needed to reconstruct shape alone. In cases where scene surface has a strong mirror-like material component, we generate highly detailed environment images, revealing room composition, objects, people, buildings, and trees visible through windows. Our approach yields state of the art view synthesis techniques, operates on low dynamic range imagery, and is robust to geometric and calibration errors. Jeong Joon Park, Aleksander Holynski, Steven M. Seitz |
CVPR | 3 |
| 2020 | Background Matting: The World Is Your Green ScreenabstractWe propose a method for creating a matte - the per-pixel foreground color and alpha - of a person by taking photos or videos in an everyday setting with a handheld camera. Most existing matting methods require a green screen background or a manually created trimap to produce a good matte. Automatic, trimap-free methods are appearing, but are not of comparable quality. In our trimap free approach, we ask the user to take an additional photo of the background without the subject at the time of capture. This step requires a small amount of foresight but is far less timeconsuming than creating a trimap. We train a deep network with an adversarial loss to predict the matte. We first train a matting network with a supervised loss on ground truth data with synthetic composites. To bridge the domain gap to real imagery with no labeling, we train another matting network guided by the first network and by a discriminator that judges the quality of composites. We demonstrate results on a wide variety of photos and videos and show significant improvement over the state of the art. Roni Sengupta, Vivek Jayaram, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
CVPR | 4 |
| 2020 | People as Scene Probes
Yifan Wang 0013, Brian Curless, Steven M. Seitz |
ECCV (10) | 3 |
| 2020 | Reconstructing NBA Players
Luyang Zhu, Konstantinos Rematas, Brian Curless, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ECCV (5) | 4 |
| 2020 | The Cone of Silence: Speech Separation by LocalizationabstractGiven a multi-microphone recording of an unknown number of speakers talking concurrently, we simultaneously localize the sources and separate the individual speakers. At the core of our method is a deep network, in the waveform domain, which isolates sources within an angular region $\theta \pm w/2$, given an angle of interest $\theta$ and angular window size $w$. By exponentially decreasing $w$, we can perform a binary search to localize and separate all sources in logarithmic time. Our algorithm also allows for an arbitrary number of potentially moving speakers at test time, including more speakers than seen during training. Experiments demonstrate state of the art performance for both source separation and source localization, particularly in high levels of background noise. Teerapat Jenrungrot, Vivek Jayaram, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
NeurIPS | 3 |
| 2018 | Surface Light Field FusionabstractWe present an approach for interactively scanning highly reflective objects with a commodity RGBD sensor. In addition to shape, our approach models the surface light field, encoding scene appearance from all directions. By factoring the surface light field into view-independent and wavelength-independent components, we arrive at a representation that can be robustly estimated with IR-equipped commodity depth sensors, and achieves high quality results. Jeong Joon Park, Richard A. Newcombe, Steven M. Seitz |
3DV | 3 |
| 2018 | Soccer on Your TabletopabstractWe present a system that transforms a monocular video of a soccer game into a moving 3D reconstruction, in which the players and field can be rendered interactively with a 3D viewer or through an Augmented Reality device. At the heart of our paper is an approach to estimate the depth map of each player, using a CNN that is trained on 3D player data extracted from soccer video games. We compare with state of the art body pose and depth estimation techniques, and show results on both synthetic ground truth benchmarks, and real YouTube soccer footage. Konstantinos Rematas, Ira Kemelmacher-Shlizerman, Brian Curless, Steven M. Seitz |
CVPR | 4 |
| 2018 | LookinGood: enhancing performance capture with real-time neural re-renderingabstractMotivated by augmented and virtual reality applications such as telepresence, there has been a recent focus in real-time performance capture of humans under motion. However, given the real-time constraint, these systems often suffer from artifacts in geometry and texture such as holes and noise in the final rendering, poor lighting, and low-resolution textures. We take the novel approach to augment such real-time performance capture systems with a deep architecture that takes a rendering from an arbitrary viewpoint, and jointly performs completion, super resolution, and denoising of the imagery in real-time. We call this approach neural (re-)rendering , and our live system "LookinGood". Our deep architecture is trained to produce high resolution and high quality images from a coarse rendering in real-time. First, we propose a self-supervised training method that does not require manual ground-truth annotation. We contribute a specialized reconstruction error that uses semantic information to focus on relevant parts of the subject, e.g. the face. We also introduce a salient reweighing scheme of the loss function that is able to discard outliers. We specifically design the system for virtual and augmented reality headsets where the consistency between the left and right eye plays a crucial role in the final user experience. Finally, we generate temporally stable results by explicitly minimizing the difference between two consecutive frames. We tested the proposed system in two different scenarios: one involving a single RGB-D sensor, and upper body reconstruction of an actor, the second consisting of full body 360° capture. Through extensive experimentation, we demonstrate how our system generalizes across unseen sequences and subjects. Ricardo Martin-Brualla, Rohit Pandey, Shuoran Yang, Pavel Pidlypenskyi, Jonathan Taylor 0001, Julien P. C. Valentin, Sameh Khamis, Philip Davidson, Anastasia Tkach, Peter Lincoln, Adarsh Kowdle, Christoph Rhemann, Dan B. Goldman, Cem Keskin, Steven M. Seitz, Shahram Izadi, Sean Ryan Fanello |
ACM Trans. Graph. | 15 |
| 2018 | PhotoShape: photorealistic materials for large-scale shape collectionsabstractExisting online 3D shape repositories contain thousands of 3D models but lack photorealistic appearance. We present an approach to automatically assign high-quality, realistic appearance models to large scale 3D shape collections. The key idea is to jointly leverage three types of online data - shape collections, material collections, and photo collections, using the photos as reference to guide assignment of materials to shapes. By generating a large number of synthetic renderings, we train a convolutional neural network to classify materials in real photos, and employ 3D-2D alignment techniques to transfer materials to different parts of each shape model. Our system produces photorealistic, relightable, 3D shapes (PhotoShapes). Keunhong Park, Konstantinos Rematas, Ali Farhadi, Steven M. Seitz |
ACM Trans. Graph. | 4 |
| 2017 | A Visual Cloud for Virtual Reality Applications
Magdalena Balazinska, Luis Ceze, Alvin Cheung, Brian Curless, Steven M. Seitz |
CIDR | 5 |
| 2017 | IM2CADabstractGiven a single photo of a room and a large database of furniture CAD models, our goal is to reconstruct a scene that is as similar as possible to the scene depicted in the photograph, and composed of objects drawn from the database. We present a completely automatic system to address this IM2CAD problem that produces high quality results on challenging imagery from interior home design and remodeling websites. Our approach iteratively optimizes the placement and scale of objects in the room to best match scene renderings to the input photo, using image comparison metrics trained via deep convolutional neural nets. By operating jointly on the full scene at once, we account for inter-object occlusions. We also show the applicability of our method in standard scene understanding benchmarks where we obtain significant improvement. Hamid Izadinia, Qi Shan, Steven M. Seitz |
CVPR | 3 |
| 2017 | Pepper's Cone: An Inexpensive Do-It-Yourself 3D DisplayabstractThis paper describes a simple 3D display that can be built from a tablet computer and a plastic sheet folded into a cone. This display allows naturally viewing a three-dimensional object from any direction over a 360-degree path of travel without the use of a head mount or special glasses. Inspired by the classic Pepper's Ghost illusion, our approach uses a curved transparent surface to reflect the image displayed on a 2D display. By properly pre-distorting the displayed image our system can produce a perspective-correct image to the viewer that appears to be suspended inside the reflector. We use the gyroscope integrated into modern tablet computers to adjust the rendered image based on the relative orientation of the viewer. The end result is a natural and intuitive interface for inspecting a 3D object. Our choice of a cone reflector is obtained by analyzing optical performance and stereo-compatibility over rotationally-symmetric conic reflector shapes. We also present the prototypes we built and measure the performance of our display through side-by-side comparisons with reference images. Jason Lawrence, Steven M. Seitz |
UIST | 3 |
| 2017 | Interactive Room Capture on 3D-Aware Mobile DevicesabstractWe propose a novel interactive system to simplify the process of indoor 3D CAD room modeling. Traditional room modeling methods require users to measure room and furniture dimensions, and manually select models that match the scene from large catalogs. Users then employ a mouse and keyboard interface to construct walls and place the objects in their appropriate locations. In contrast, our system leverages the sensing capabilities of a 3D aware mobile device, recent advances in object recognition, and a novel augmented reality user interface, to capture indoor 3D room models in-situ. With a few taps, a user can mark the surface of an object, take a photo, and the system retrieves and places a matching 3D model into the scene, from a large online database. User studies indicate that this modality is significantly quicker, more accurate, and requires less effort than traditional desktop tools. Aditya Sankar, Steven M. Seitz |
UIST | 2 |
| 2017 | 3D Time-Lapse Reconstruction from Internet Photos
Ricardo Martin-Brualla, David Gallup, Steven M. Seitz |
Int. J. Comput. Vis. | 3 |
| 2017 | Summarizing Unconstrained Videos Using Salient MontagesabstractWe present a novel method to summarize unconstrained videos using salient montages (i.e., a "melange" of frames in the video as shown in Fig. 1, by finding "montageable moments" and identifying the salient people and actions to depict in each montage. Our method aims at addressing the increasing need for generating concise visualizations from the large number of videos being captured from portable devices. Our main contributions are (1) the process of finding salient people and moments to form a montage, and (2) the application of this method to videos taken "in the wild" where the camera moves freely. As such, we demonstrate results on head-mounted cameras, where the camera moves constantly, as well as on videos downloaded from YouTube. In our experiments, we show that our method can reliably detect and track humans under significant action and camera motion. Moreover, the predicted salient people are more accurate than results from state-of-the-art video salieny method [1] . Finally, we demonstrate that a novel "montageability" score can be used to retrieve results with relatively high precision which allows us to present high quality montages to users.We present a novel method to summarize unconstrained videos using salient montages (i.e., a "melange" of frames in the video as shown in Fig. 1, by finding "montageable moments" and identifying the salient people and actions to depict in each montage. Our method aims at addressing the increasing need for generating concise visualizations from the large number of videos being captured from portable devices. Our main contributions are (1) the process of finding salient people and moments to form a montage, and (2) the application of this method to videos taken "in the wild" where the camera moves freely. As such, we demonstrate results on head-mounted cameras, where the camera moves constantly, as well as on videos downloaded from YouTube. In our experiments, we show that our method can reliably detect and track humans under significant action and camera motion. Moreover, the predicted salient people are more accurate than results from state-of-the-art video salieny method [1] . Finally, we demonstrate that a novel "montageability" score can be used to retrieve results with relatively high precision which allows us to present high quality montages to users. Min Sun 0001, Ali Farhadi, Ben Taskar, Steven M. Seitz |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2017 | Synthesizing Obama: learning lip sync from audioabstractGiven audio of President Barack Obama, we synthesize a high quality video of him speaking with accurate lip sync, composited into a target video clip. Trained on many hours of his weekly address footage, a recurrent neural network learns the mapping from raw audio features to mouth shapes. Given the mouth shape at each time instant, we synthesize high quality mouth texture, and composite it with proper 3D pose matching to change what he appears to be saying in a target video to match the input audio track. Our approach produces photorealistic results. Supasorn Suwajanakorn, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ACM Trans. Graph. | 2 |
| 2016 | The MegaFace Benchmark: 1 Million Faces for Recognition at ScaleabstractRecent face recognition experiments on a major benchmark (LFW [15]) show stunning performance-a number of algorithms achieve near to perfect score, surpassing human recognition rates. In this paper, we advocate evaluations at the million scale (LFW includes only 13K photos of 5K people). To this end, we have assembled the MegaFace dataset and created the first MegaFace challenge. Our dataset includes One Million photos that capture more than 690K different individuals. The challenge evaluates performance of algorithms with increasing numbers of "distractors" (going from 10 to 1M) in the gallery set. We present both identification and verification performance, evaluate performance with respect to pose and a persons age, and compare as a function of training data size (#photos and #people). We report results of state of the art and baseline algorithms. The MegaFace dataset, baseline code, and evaluation scripts, are all publicly released for further experimentations1. Ira Kemelmacher-Shlizerman, Steven M. Seitz, Evan Brossard |
CVPR | 2 |
| 2016 | In situ CAD captureabstractWe present an interactive system to capture CAD-like 3D models of indoor scenes, on a mobile device. To overcome sensory and computational limitations of the mobile platform, we employ an in situ, semi-automated approach and harness the user's high-level knowledge of the scene to assist the reconstruction and modeling algorithms. The modeling proceeds in two stages: (1) The user captures the 3D shape and dimensions of the room. (2) The user then uses voice commands and an augmented reality sketching interface to insert objects of interest, such as furniture, artwork, doors and windows. Our system recognizes the sketches and add a corresponding 3D model into the scene at the appropriate location. The key contributions of this work are the design of a multi-modal user interface to effectively capture the user's semantic understanding of the scene and the underlying algorithms that process the input to produce useful reconstructions. Aditya Sankar, Steven M. Seitz |
MobileHCI | 2 |
| 2016 | Ranking Highlights in Personal Videos by Analyzing Edited VideosabstractWe present a fully automatic system for ranking domain-specific highlights in unconstrained personal videos by analyzing online edited videos. A novel latent linear ranking model is proposed to handle noisy training data harvested online. Specifically, given a targeted domain such as "surfing," our system mines the YouTube database to find pairs of raw and their corresponding edited videos. Leveraging the assumption that an edited video is more likely to contain highlights than the trimmed parts of the raw video, we obtain pair-wise ranking constraints to train our model. The learning task is challenging due to the amount of noise and variation in the mined data. Hence, a latent loss function is incorporated to mitigate the issues caused by the noise. We efficiently learn the latent model on a large number of videos (about 870 min in total) using a novel EM-like procedure. Our latent ranking model outperforms its classification counterpart and is fairly competitive compared with a fully supervised ranking system that requires labels from Amazon Mechanical Turk. We further show that a state-of-the-art audio feature mel-frequency cepstral coefficients is inferior to a state-of-the-art visual feature. By combining both audio-visual features, we obtain the best performance in dog activity, surfing, skating, and viral video domains. Finally, we show that impressive highlights can be detected without additional human supervision for seven domains (i.e., skating, surfing, skiing, gymnastics, parkour, dog activity, and viral video) in unconstrained personal videos. Min Sun 0001, Ali Farhadi, Tseng-Hung Chen, Steven M. Seitz |
IEEE Trans. Image Process. | 4 |
| 2016 | Jump: virtual reality videoabstractWe present Jump, a practical system for capturing high resolution, omnidirectional stereo (ODS) video suitable for wide scale consumption in currently available virtual reality (VR) headsets. Our system consists of a video camera built using off-the-shelf components and a fully automatic stitching pipeline capable of capturing video content in the ODS format. We have discovered and analyzed the distortions inherent to ODS when used for VR display as well as those introduced by our capture method and show that they are small enough to make this approach suitable for capturing a wide variety of scenes. Our stitching algorithm produces robust results by reducing the problem to one of pairwise image interpolation followed by compositing. We introduce novel optical flow and compositing methods designed specifically for this task. Our algorithm is temporally coherent and efficient, is currently running at scale on a distributed computing platform, and is capable of processing hours of footage each day. David Gallup, Jonathan T. Barron, Janne Kontkanen, Noah Snavely, Sameer Agarwal 0001, Steven M. Seitz |
ACM Trans. Graph. | 8 |
| 2015 | DynamicFusion: Reconstruction and tracking of non-rigid scenes in real-timeabstractWe present the first dense SLAM system capable of reconstructing non-rigidly deforming scenes in real-time, by fusing together RGBD scans captured from commodity sensors. Our DynamicFusion approach reconstructs scene geometry whilst simultaneously estimating a dense volumetric 6D motion field that warps the estimated geometry into a live frame. Like KinectFusion, our system produces increasingly denoised, detailed, and complete reconstructions as more measurements are fused, and displays the updated model in real time. Because we do not require a template or other prior scene model, the approach is applicable to a wide range of moving objects and scenes. Richard A. Newcombe, Dieter Fox, Steven M. Seitz |
CVPR | 3 |
| 2015 | Depth from focus with your mobile phoneabstractWhile prior depth from focus and defocus techniques operated on laboratory scenes, we introduce the first depth from focus (DfF) method capable of handling images from mobile phones and other hand-held cameras. Achieving this goal requires solving a novel uncalibrated DfF problem and aligning the frames to account for scene parallax. Our approach is demonstrated on a range of challenging cases and produces high quality results. Supasorn Suwajanakorn, Steven M. Seitz |
CVPR | 3 |
| 2015 | 3D Time-Lapse Reconstruction from Internet PhotosabstractGiven an Internet photo collection of a landmark, we compute a 3D time-lapse video sequence where a virtual camera moves continuously in time and space. While previous work assumed a static camera, the addition of camera motion during the time-lapse creates a very compelling impression of parallax. Achieving this goal, however, requires addressing multiple technical challenges, including solving for time-varying depth maps, regularizing 3D point color profiles over time, and reconstructing high quality, hole-free images at every frame from the projected profiles. Our results show photorealistic time-lapses of skylines and natural scenes over many years, with dramatic parallax effects. Ricardo Martin-Brualla, David Gallup, Steven M. Seitz |
ICCV | 3 |
| 2015 | What Makes Tom Hanks Look Like Tom HanksabstractWe reconstruct a controllable model of a person from a large photo collection that captures his or her persona, i.e., physical appearance and behavior. The ability to operate on unstructured photo collections enables modeling a huge number of people, including celebrities and other well photographed people without requiring them to be scanned. Moreover, we show the ability to drive or puppeteer the captured person B using any other video of a different person A. In this scenario, B acts out the role of person A, but retains his/her own personality and character. Our system is based on a novel combination of 3D face reconstruction, tracking, alignment, and multi-texture modeling, applied to the puppeteering problem. We demonstrate convincing results on a large variety of celebrities derived from Internet imagery and video. Supasorn Suwajanakorn, Steven M. Seitz, Ira Kemelmacher-Shlizerman |
ICCV | 2 |
| 2015 | Time-lapse mining from internet photosabstractWe introduce an approach for synthesizing time-lapse videos of popular landmarks from large community photo collections. The approach is completely automated and leverages the vast quantity of photos available online. First, we cluster 86 million photos into landmarks and popular viewpoints. Then, we sort the photos by date and warp each photo onto a common viewpoint. Finally, we stabilize the appearance of the sequence to compensate for lighting effects and minimize flicker. Our resulting time-lapses show diverse changes in the world's most popular sites, like glaciers shrinking, skyscrapers being constructed, and waterfalls changing course. Ricardo Martin-Brualla, David Gallup, Steven M. Seitz |
ACM Trans. Graph. | 3 |
| 2014 | Accurate Geo-Registration by Ground-to-Aerial Image MatchingabstractWe address the problem of geo-registering ground-based multi-view stereo models by ground-to-aerial image matching. The main contribution is a fully automated geo-registration pipeline with a novel viewpoint-dependent matching method that handles ground to aerial viewpoint variation. We conduct large-scale experiments which consist of many popular outdoor landmarks in Rome. The proposed approach demonstrates a high success rate for the task, and dramatically outperforms state-of-the-art techniques, yielding geo-registration at pixel-level accuracy. Qi Shan, Changchang Wu, Brian Curless, Yasutaka Furukawa, Steven M. Seitz |
3DV | 6 |
| 2014 | Illumination-Aware Age ProgressionabstractWe present an approach that takes a single photograph of a child as input and automatically produces a series of age-progressed outputs between 1 and 80 years of age, accounting for pose, expression, and illumination. Leveraging thousands of photos of children and adults at many ages from the Internet, we first show how to compute average image subspaces that are pixel-to-pixel aligned and model variable lighting. These averages depict a prototype man and woman aging from 0 to 80, under any desired illumination, and capture the differences in shape and texture between ages. Applying these differences to a new photo yields an age progressed result. Contributions include relightable age subspaces, a novel technique for subspace-to-subspace alignment, and the most extensive evaluation of age progression techniques in the literature. Ira Kemelmacher-Shlizerman, Supasorn Suwajanakorn, Steven M. Seitz |
CVPR | 3 |
| 2014 | Occluding Contours for Multi-view StereoabstractThis paper leverages occluding contours (aka "internal silhouettes") to improve the performance of multi-view stereo methods. The contributions are 1) a new technique to identify free-space regions arising from occluding contours, and 2) a new approach for incorporating the resulting free-space constraints into Poisson surface reconstruction. The proposed approach outperforms state of the art MVS techniques for challenging Internet datasets, yielding dramatic quality improvements both around object contours and in surface detail. Qi Shan, Brian Curless, Yasutaka Furukawa, Steven M. Seitz |
CVPR | 5 |
| 2014 | The 3D Jigsaw Puzzle: Mapping Large Indoor Spaces
Ricardo Martin-Brualla, Yanling He, Bryan C. Russell, Steven M. Seitz |
ECCV (3) | 4 |
| 2014 | Photo Uncrop
Qi Shan, Brian Curless, Yasutaka Furukawa, Steven M. Seitz |
ECCV (6) | 5 |
| 2014 | Ranking Domain-Specific Highlights by Analyzing Edited Videos
Min Sun 0001, Ali Farhadi, Steven M. Seitz |
ECCV (1) | 3 |
| 2014 | Salient Montages from Unconstrained Videos
Min Sun 0001, Ali Farhadi, Ben Taskar, Steven M. Seitz |
ECCV (7) | 4 |
| 2014 | Total Moving Face Reconstruction
Supasorn Suwajanakorn, Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ECCV (4) | 3 |
| 2013 | Single View Reconstruction of Piecewise Swept SurfacesabstractWe present a novel approach for the single view reconstruction (SVR) of piecewise swept scenes, exploiting the regular structure present in these man-made scenes. The parallelism of lines and its extension to curves are used as cues within a novel sequential algorithm propagating information across faces and their boundaries. Using this approach we are able to model a wide variety of architectural scenes and man-made objects. We show results generated both automatically as well as using user interaction, within a unified framework. Avanish Kushal, Steven M. Seitz |
3DV | 2 |
| 2013 | The Visual Turing Test for Scene ReconstructionabstractWe present the first large scale system for capturing and rendering relight able scene reconstructions from massive unstructured photo collections taken under different illumination conditions and viewpoints. We combine photos taken from many sources, Flickr-Based ground-level imagery, oblique aerial views, and street view, to recover models that are significantly more complete and detailed than previously demonstrated. We demonstrate the ability to match both the viewpoint and illumination of arbitrary input photos, enabling a Visual Turing Test in which photo and rendering are viewed side-by-side and the observer has to guess which is which. While we cannot yet fool human perception, the gap is closing. Qi Shan, Riley Adams, Brian Curless, Yasutaka Furukawa, Steven M. Seitz |
3DV | 5 |
| 2013 | 3D Wikipedia: using online text to automatically label and navigate reconstructed geometryabstractWe introduce an approach for analyzing Wikipedia and other text, together with online photos, to produce annotated 3D models of famous tourist sites. The approach is completely automated, and leverages online text and photo co-occurrences via Google Image Search. It enables a number of new interactions, which we demonstrate in a new 3D visualization tool. Text can be selected to move the camera to the corresponding objects, 3D bounding boxes provide anchors back to the text describing them, and the overall narrative of the text provides a temporal guide for automatically flying through the scene to visualize the world as you read about it. We show compelling results on several major tourist sites. Bryan C. Russell, Ricardo Martin-Brualla, Daniel J. Butler, Steven M. Seitz, Luke Zettlemoyer |
ACM Trans. Graph. | 4 |
| 2013 | Navigating the worldwide community of photosabstractThe last decade has seen an explosion in the number of photographs available on the Internet. The sheer volume of interesting photos makes it a challenge to explore this space. Various Web and social media sites, along with search and indexing techniques, have been developed in response. One natural way to navigate these images in a 3D geo-located context. In this article, we reflect on our work in this area, with a focus on techniques that build partial 3D scene models to help find and navigate interesting photographs in an interactive, immersive 3D setting. We also discuss how finding such relationships among photographs opens up exciting new possibilities for multimedia authoring, visualization, and editing. Richard Szeliski, Noah Snavely, Steven M. Seitz |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Collection flowabstractComputing optical flow between any pair of Internet face photos is challenging for most current state of the art flow estimation methods due to differences in illumination, pose, and geometry. We show that flow estimation can be dramatically improved by leveraging a large photo collection of the same (or similar) object. In particular, consider the case of photos of a celebrity from Google Image Search. Any two such photos may have different facial expression, lighting and face orientation. The key idea is that instead of computing flow directly between the input pair (I, J), we compute versions of the images (I', J') in which facial expressions and pose are normalized while lighting is preserved. This is achieved by iteratively projecting each photo onto an appearance subspace formed from the full photo collection. The desired flow is obtained through concatenation of flows (I → I') o (J' → J). Our approach can be used with any two-frame optical flow algorithm, and significantly boosts the performance of the algorithm by providing invariance to lighting and shape changes. Ira Kemelmacher-Shlizerman, Steven M. Seitz |
CVPR | 2 |
| 2012 | Schematic surface reconstructionabstractThis paper introduces a schematic representation for architectural scenes together with robust algorithms for reconstruction from sparse 3D point cloud data. The schematic models architecture as a network of transport curves, approximating a floorplan, with associated profile curves, together comprising an interconnected set of swept surfaces. The representation is extremely concise, composed of a handful of planar curves, and easily interpretable by humans. The approach also provides a principled mechanism for interpolating a dense surface, and enables filling in holes in the data, by means of a pipeline that employs a global optimization over all parameters. By incorporating a displacement map on top of the schematic surface, it is possible to recover fine details. Experiments show the ability to reconstruct extremely clean and simple models from sparse structure-from-motion point clouds of complex architectural scenes. Changchang Wu, Sameer Agarwal 0001, Brian Curless, Steven M. Seitz |
CVPR | 4 |
| 2012 | Capturing indoor scenes with smartphonesabstractIn this paper, we present a novel smartphone application designed to easily capture, visualize and reconstruct homes, offices and other indoor scenes. Our application leverages data from smartphone sensors such as the camera, accelerometer, gyroscope and magnetometer to help model the indoor scene. The output of the system is two-fold; first, an interactive visual tour of the scene is generated in real time that allows the user to explore each room and transition between connected rooms. Second, with some basic interactive photogrammetric modeling the system generates a 2D floor plan and accompanying 3D model of the scene, under a Manhattan-world assumption. The approach does not require any specialized equipment or training and is able to produce accurate floor plans. Aditya Sankar, Steven M. Seitz |
UIST | 2 |
| 2011 | Binocular Photometric StereoabstractThis paper considers the problem of computing scene depth from a stereo pair of cameras under a sequence of illumination directions. By integrating parallax and shading cues, we obtain both metric depth and fine surface details. Casting this problem into the filter flow framework [16], enables a convex formulation of the problem, and thus a globally optimal solution. We demonstrate high quality, continuous depth maps on a range of examples. Hao Du 0004, Dan B. Goldman, Steven M. Seitz |
BMVC | 3 |
| 2011 | Where's Waldo: Matching people in images of crowdsabstractGiven a community-contributed set of photos of a crowded public event, this paper addresses the problem of finding all images of each person in the scene. This problem is very challenging due to large changes in camera viewpoints, severe occlusions, low resolution and photos from tens or hundreds of different photographers. Despite these challenges, the problem is made tractable by exploiting a variety of visual and contextual cues-appearance, time-stamps, camera pose and co-occurrence of people. This paper demonstrates an approach that integrates these cues to enable high quality person matching in community photo collections downloaded from Flickr.com. Rahul Garg 0002, Steven M. Seitz, Deva Ramanan, Noah Snavely |
CVPR | 2 |
| 2011 | Multicore bundle adjustmentabstractWe present the design and implementation of new inexact Newton type Bundle Adjustment algorithms that exploit hardware parallelism for efficiently solving large scale 3D scene reconstruction problems. We explore the use of multicore CPU as well as multicore GPUs for this purpose. We show that overcoming the severe memory and bandwidth limitations of current generation GPUs not only leads to more space efficient algorithms, but also to surprising savings in runtime. Our CPU based system is up to ten times and our GPU based system is up to thirty times faster than the current state of the art methods, while maintaining comparable convergence behavior. The code and additional results are available at http://grail.cs.washington.edu/projects/mcba. Changchang Wu, Sameer Agarwal 0001, Brian Curless, Steven M. Seitz |
CVPR | 4 |
| 2011 | Interactive 3D modeling of indoor environments with a consumer depth cameraabstractDetailed 3D visual models of indoor spaces, from walls and floors to objects and their configurations, can provide extensive knowledge about the environments as well as rich contextual information of people living therein. Vision-based 3D modeling has only seen limited success in applications, as it faces many technical challenges that only a few experts understand, let alone solve. In this work we utilize (Kinect style) consumer depth cameras to enable non-expert users to scan their personal spaces into 3D models. We build a prototype mobile system for 3D modeling that runs in real-time on a laptop, assisting and interacting with the user on-the-fly. Color and depth are jointly used to achieve robust 3D registration. The system offers online feedback and hints, tolerates human errors and alignment failures, and helps to obtain complete scene coverage. We show that our prototype system can both scan large environments (50 meters across) and at the same time preserve fine details (centimeter accuracy). The capability of detailed 3D modeling leads to many promising applications such as accurate 3D localization, measuring dimensions, and interactive visualization. Hao Du 0004, Peter Henry, Xiaofeng Ren, Marvin Cheng, Dan B. Goldman, Steven M. Seitz, Dieter Fox |
UbiComp | 6 |
| 2011 | Face reconstruction in the wildabstractWe address the problem of reconstructing 3D face models from large unstructured photo collections, e.g., obtained by Google image search or from personal photo collections in iPhoto. This problem is extremely challenging due to the high degree of variability in pose, illumination, facial expression, non-rigid changes in face shape and reflectance over time and occlusions. In light of this extreme variability, no single reconstruction can be consistent with all of the images. Instead, we define as the goal of reconstruction to recover a model that is locally consistent with the image set. I.e., each local region of the model is consistent with a large set of photos, resulting in a model that captures the dominant trends in the input data for different parts of the face. Our approach leverages multi-image shading, but unlike traditional photometric stereo approaches, allows for changes in viewpoint and shape. We optimize over pose, shape, and lighting in an iterative approach that seeks to minimize the rank of the transformed images. This approach produces high quality shape models for a wide range of celebrities from photos available on the Internet. Ira Kemelmacher-Shlizerman, Steven M. Seitz |
ICCV | 2 |
| 2011 | Exploring photobiosabstractWe present an approach for generating face animations from large image collections of the same person. Such collections, which we call photobios , sample the appearance of a person over changes in pose, facial expression, hairstyle, age, and other variations. By optimizing the order in which images are displayed and cross-dissolving between them, we control the motion through face space and create compelling animations (e.g., render a smooth transition from frowning to smiling). Used in this context, the cross dissolve produces a very strong motion effect; a key contribution of the paper is to explain this effect and analyze its operating range. The approach operates by creating a graph with faces as nodes, and similarities as edges, and solving for walks and shortest paths on this graph. The processing pipeline involves face detection, locating fiducials (eyes/nose/mouth), solving for pose, warping to frontal views, and image comparison based on Local Binary Patterns. We demonstrate results on a variety of datasets including time-lapse photography, personal photo collections, and images of celebrities downloaded from the Internet. Our approach is the basis for the Face Movies feature in Google's Picasa. Ira Kemelmacher-Shlizerman, Eli Shechtman, Rahul Garg 0002, Steven M. Seitz |
ACM Trans. Graph. | 4 |
| 2010 | Towards Internet-scale multi-view stereoabstractThis paper introduces an approach for enabling existing multi-view stereo methods to operate on extremely large unstructured photo collections. The main idea is to decompose the collection into a set of overlapping sets of photos that can be processed in parallel, and to merge the resulting reconstructions. This overlapping clustering problem is formulated as a constrained optimization and solved iteratively. The merging algorithm, designed to be parallel and out-of-core, incorporates robust filtering steps to eliminate low-quality reconstructions and enforce global visibility constraints. The approach has been tested on several large datasets downloaded from Flickr.com, including one with over ten thousand images, yielding a 3D reconstruction with nearly thirty million points. Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski |
CVPR | 3 |
| 2010 | Generating sharp panoramas from motion-blurred videosabstractIn this paper, we show how to generate a sharp panorama from a set of motion-blurred video frames. Our technique is based on joint global motion estimation and multi-frame deblurring. It also automatically computes the duty cycle of the video, namely the percentage of time between frames that is actually exposure time. The duty cycle is necessary for allowing the blur kernels to be accurately extracted and then removed. We demonstrate our technique on a number of videos. Yunpeng Li 0002, Sing Bing Kang, Neel Joshi, Steven M. Seitz, Daniel P. Huttenlocher |
CVPR | 4 |
| 2010 | Regenerative morphingabstractWe present a new image morphing approach in which the output sequence is regenerated from small pieces of the two source (input) images. The approach does not require manual correspondence, and generates compelling results even when the images are of very different objects (e.g., a cloud and a face). We pose the morphing task as an optimization with the objective of achieving bidirectional similarity of each frame to its neighbors, and also to the source images. The advantages of this approach are 1) it can operate fully automatically, producing effective results for many sequences (but also supports manual correspondences, when available), 2) ghosting artifacts are minimized, and 3) different parts of the scene move at different rates, yielding more interesting (and less robotic) transitions. Eli Shechtman, Alex Rav-Acha, Michal Irani, Steven M. Seitz |
CVPR | 4 |
| 2010 | Bundle Adjustment in the Large
Sameer Agarwal 0001, Noah Snavely, Steven M. Seitz, Richard Szeliski |
ECCV (2) | 3 |
| 2010 | Being John Malkovich
Ira Kemelmacher-Shlizerman, Aditya Sankar, Eli Shechtman, Steven M. Seitz |
ECCV (1) | 4 |
| 2010 | Shape and Spatially-Varying BRDFs from Photometric StereoabstractThis paper describes a photometric stereo method designed for surfaces with spatially-varying BRDFs, including surfaces with both varying diffuse and specular properties. Our optimization-based method builds on the observation that most objects are composed of a small number of fundamental materials by constraining each pixel to be representable by a combination of at most two such materials. This approach recovers not only the shape but also material BRDFs and weight maps, yielding accurate rerenderings under novel lighting conditions for a wide variety of objects. We demonstrate examples of interactive editing operations made possible by our approach. Dan B. Goldman, Brian Curless, Aaron Hertzmann, Steven M. Seitz |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2010 | Scene Reconstruction and Visualization From Community Photo CollectionsabstractThere are billions of photographs on the Internet, representing an extremely large, rich, and nearly comprehensive visual record of virtually every famous place on Earth. Unfortunately, these massive community photo collections are almost completely unstructured, making it very difficult to use them for applications such as the virtual exploration of our world. Over the past several years, advances in computer vision have made it possible to automatically reconstruct 3-D geometry - including camera positions and scene models - from these large, diverse photo collections. Once the geometry is known, we can recover higher level information from the spatial distribution of photos, such as the most common viewpoints and paths through the scene. This paper reviews recent progress on these challenging computer vision problems, and describes how we can use the recovered structure to turn community photo collections into immersive, interactive 3-D experiences. Noah Snavely, Ian Simon, Michael Goesele, Richard Szeliski, Steven M. Seitz |
Proc. IEEE | 5 |
| 2009 | Manhattan-world stereoabstractMulti-view stereo (MVS) algorithms now produce reconstructions that rival laser range scanner accuracy. However, stereo algorithms require textured surfaces, and therefore work poorly for many architectural scenes (e.g., building interiors with textureless, painted walls). This paper presents a novel MVS approach to overcome these limitations for Manhattan World scenes, i.e., scenes that consists of piece-wise planar surfaces with dominant directions. Given a set of calibrated photographs, we first reconstruct textured regions using an existing MVS algorithm, then extract dominant plane directions, generate plane hypotheses, and recover per-view depth maps using Markov random fields. We have tested our algorithm on several datasets ranging from office interiors to outdoor buildings, and demonstrate results that outperform the current state of the art for such texture-poor scenes. Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski |
CVPR | 3 |
| 2009 | Building Rome in a dayabstractWe present a system that can match and reconstruct 3D scenes from extremely large collections of photographs such as those found by searching for a given city (e.g., Rome) on Internet photo sharing sites. Our system uses a collection of novel parallel distributed matching and reconstruction algorithms, designed to maximize parallelism at each stage in the pipeline and minimize serialization bottlenecks. It is designed to scale gracefully with both the size of the problem and the amount of available computation. We have experimented with a variety of alternative algorithms at each stage of the pipeline and report on which ones work best in a parallel computing environment. Our experimental results demonstrate that it is now possible to reconstruct cities consisting of 150 K images in less than a day on a cluster with 500 compute cores. Sameer Agarwal 0001, Noah Snavely, Ian Simon, Steven M. Seitz, Richard Szeliski |
ICCV | 4 |
| 2009 | Reconstructing building interiors from imagesabstractThis paper proposes a fully automated 3D reconstruction and visualization system for architectural scenes (interiors and exteriors). The reconstruction of indoor environments from photographs is particularly challenging due to texture-poor planar surfaces such as uniformly-painted walls. Our system first uses structure-from-motion, multi-view stereo, and a stereo algorithm specifically designed for Manhattan-world scenes (scenes consisting predominantly of piece-wise planar surfaces with dominant directions) to calibrate the cameras and to recover initial 3D geometry in the form of oriented points and depth maps. Next, the initial geometry is fused into a 3D model with a novel depth-map integration algorithm that, again, makes use of Manhattan-world assumptions and produces simplified 3D models. Finally, the system enables the exploration of reconstructed environments with an interactive, image-based 3D viewer. We demonstrate results on several challenging datasets, including a 3D reconstruction and image-based walk-through of an entire floor of a house, the first result of this kind from an automated computer vision system. Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski |
ICCV | 3 |
| 2009 | The dimensionality of scene appearanceabstractLow-rank approximation of image collections (e.g., via PCA) is a popular tool in many areas of computer vision. Yet, surprisingly little is known justifying the observation that images of an object or scene tend to be low dimensional, beyond the special case of Lambertian scenes. This paper considers the question of how many basis images are needed to span the space of images of a scene under real-world lighting and viewing conditions, allowing for general BRDFs. We establish new theoretical upper bounds on the number of basis images necessary to represent a wide variety of scenes under very general conditions, and perform empirical studies to justify the assumptions. We then demonstrate a number of novel applications of linear models for scene appearance for Internet photo collections. These applications include, image reconstruction, occluder-removal, and expanding field of view. Rahul Garg 0002, Hao Du 0004, Steven M. Seitz, Noah Snavely |
ICCV | 3 |
| 2009 | Filter flowabstractThe filter flow problem is to compute a space-variant linear filter that transforms one image into another. This framework encompasses a broad range of transformations including stereo, optical flow, lighting changes, blur, and combinations of these effects. Parametric models such as affine motion, vignetting, and radial distortion can also be modeled within the same framework. All such transformations are modeled by selecting a number of constraints and objectives on the filter entries from a catalog which we enumerate. Most of the constraints are linear, leading to globally optimal solutions (via linear programming) for affine transformations, depth-from-defocus, and other problems. Adding a (non-convex) compactness objective enables solutions for optical flow with illumination changes, space-variant defocus, and higher-order smoothness. Steven M. Seitz, Simon Baker |
ICCV | 1 |
| 2009 | Rectified Surface MosaicsabstractWe approach mosaicing as a camera tracking problem within a known parameterized surface. From a video of a camera moving within a surface, we compute a mosaic representing the texture of that surface, flattened onto a planar image. Our approach works by defining a warp between images as a function of surface geometry and camera pose. Globally optimizing this warp to maximize alignment across all frames determines the camera trajectory, and the corresponding flattened mosaic image. In contrast to previous mosaicing methods which assume planar or distant scenes, or controlled camera motion, our approach enables mosaicing in cases where the camera moves unpredictably through proximal surfaces, such as in medical endoscopy applications. Robert E. Carroll, Steven M. Seitz |
Int. J. Comput. Vis. | 2 |
| 2008 | Fast algorithms for L∞ problems in multiview geometryabstractMany problems in multi-view geometry, when posed as minimization of the maximum reprojection error across observations, can be solved optimally in polynomial time. We show that these problems are instances of a convex-concave generalized fractional program. We survey the major solution methods for solving problems of this form and present them in a unified framework centered around a single parametric optimization problem. We propose two new algorithms and show that the algorithm proposed by Olsson et al. [21] is a special case of a classical algorithm for generalized fractional programming. The performance of all the algorithms is compared on a variety of datasets, and the algorithm proposed by Gugat [12] stands out as a clear winner. An open source MATLAB toolbox that implements all the algorithms presented here is made available. Sameer Agarwal 0001, Noah Snavely, Steven M. Seitz |
CVPR | 3 |
| 2008 | Skeletal graphs for efficient structure from motionabstractWe address the problem of efficient structure from motion for large, unordered, highly redundant, and irregularly sampled photo collections, such as those found on Internet photo-sharing sites. Our approach computes a small skeletal subset of images, reconstructs the skeletal set, and adds the remaining images using pose estimation. Our technique drastically reduces the number of parameters that are considered, resulting in dramatic speedups, while provably approximating the covariance of the full set of parameters. To compute a skeletal image set, we first estimate the accuracy of two-frame reconstructions between pairs of overlapping images, then use a graph algorithm to select a subset of images that, when reconstructed, approximates the accuracy of the full set. A final bundle adjustment can then optionally be used to restore any loss of accuracy. Noah Snavely, Steven M. Seitz, Richard Szeliski |
CVPR | 2 |
| 2008 | Scene Segmentation Using the Wisdom of Crowds
Ian Simon, Steven M. Seitz |
ECCV (2) | 2 |
| 2008 | Video object annotation, navigation, and compositionabstractWe explore the use of tracked 2D object motion to enable novel approaches to interacting with video. These include moving annotations, video navigation by direct manipulation of objects, and creating an image composite from multiple video frames. Features in the video are automatically tracked and grouped in an off-line preprocess that enables later interactive manipulation. Examples of annotations include speech and thought balloons, video graffiti, path arrows, video hyperlinks, and schematic storyboards. We also demonstrate a direct-manipulation interface for random frame access using spatial constraints, and a drag-and-drop interface for assembling still images from videos. Taken together, our tools can be employed in a variety of applications including film and video editing, visual tagging, and authoring rich media such as hyperlinked video. Dan B. Goldman, Chris Gonterman, Brian Curless, David Salesin, Steven M. Seitz |
UIST | 5 |
| 2008 | Modeling the World from Internet Photo Collections
Noah Snavely, Steven M. Seitz, Richard Szeliski |
Int. J. Comput. Vis. | 2 |
| 2008 | Reconstructing relief surfaces
George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla |
Image Vis. Comput. | 3 |
| 2008 | Finding paths through the world's photosabstractWhen a scene is photographed many times by different people, the viewpoints often cluster along certain paths. These paths are largely specific to the scene being photographed, and follow interesting regions and viewpoints. We seek to discover a range of such paths and turn them into controls for image-based rendering. Our approach takes as input a large set of community or personal photos, reconstructs camera viewpoints, and automatically computes orbits, panoramas, canonical views, and optimal paths between views. The scene can then be interactively browsed in 3D using these controls or with six degree-of-freedom free-viewpoint control. As the user browses the scene, nearby views are continuously selected and transformed, using control-adaptive reprojection techniques. Noah Snavely, Rahul Garg 0002, Steven M. Seitz, Richard Szeliski |
ACM Trans. Graph. | 3 |
| 2007 | A Probabilistic Model for Object Recognition, Segmentation, and Non-Rigid CorrespondenceabstractWe describe a method for fully automatic object recognition and segmentation using a set of reference images to specify the appearance of each object. Our method uses a generative model of image formation that takes into account occlusions, simple lighting changes, and object deformations. We take advantage of local features to identify, locate, and extract multiple objects in the presence of large viewpoint changes, nonrigid motions with large numbers of degrees of freedom, occlusions, and clutter. We simultaneously compute an object-level segmentation and a dense correspondence between the pixels of the appropriate reference images and the image to be segmented. Ian Simon, Steven M. Seitz |
CVPR | 2 |
| 2007 | Rectified Surface MosaicsabstractWe approach mosaicing as a camera tracking problem within a known parameterized surface. From a video of a camera moving within a surface, we compute a mosaic representing the texture of that surface, flattened onto a planar image. Our approach works by defining a warp between images as a function of surface geometry and camera pose. Globally optimizing this warp to maximize alignment across all frames determines the camera trajectory, and the corresponding flattened mosaic image. In contrast to previous mosaicing methods which assume planar or distant scenes, or controlled camera motion, our approach enables mosaicing in cases where the camera moves unpredictably through proximal surfaces, such as in medical endoscopy applications. Robert E. Carroll, Steven M. Seitz |
ICCV | 2 |
| 2007 | Multi-View Stereo for Community Photo CollectionsabstractWe present a multi-view stereo algorithm that addresses the extreme changes in lighting, scale, clutter, and other effects in large online community photo collections. Our idea is to intelligently choose images to match, both at a per-view and per-pixel level. We show that such adaptive view selection enables robust performance even with dramatic appearance variability. The stereo matching technique takes as input sparse 3D points reconstructed from structure-from-motion methods and iteratively grows surfaces from these points. Optimizing for surface normals within a photoconsistency measure significantly improves the matching results. While the focus of our approach is to estimate high-quality depth maps, we also show examples of merging the resulting depth maps into compelling scene reconstructions. We demonstrate our algorithm on standard multi-view stereo datasets and on casually acquired photo collections of famous scenes gathered from the Internet. Michael Goesele, Noah Snavely, Brian Curless, Hugues Hoppe, Steven M. Seitz |
ICCV | 5 |
| 2007 | Scene Summarization for Online Image CollectionsabstractWe formulate the problem of scene summarization as selecting a set of images that efficiently represents the visual content of a given scene. The ideal summary presents the most interesting and important aspects of the scene with minimal redundancy. We propose a solution to this problem using multi-user image collections from the Internet. Our solution examines the distribution of images in the collection to select a set of canonical views to form the scene summary, using clustering techniques on visual features. The summaries we compute also lend themselves naturally to the browsing of image collections, and can be augmented by analyzing user-specified image tag data. We demonstrate the approach using a collection of images of the city of Rome, showing the ability to automatically decompose the images into separate scenes, and identify canonical views for each scene. Ian Simon, Noah Snavely, Steven M. Seitz |
ICCV | 3 |
| 2007 | Estimating Optimal Parameters for MRF Stereo from a Single Image PairabstractThis paper presents a novel approach for estimating the parameters for MRF-based stereo algorithms. This approach is based on a new formulation of stereo as a maximum a posterior (MAP) problem in which both a disparity map and MRF parameters are estimated from the stereo pair itself. We present an iterative algorithm for the MAP estimation that alternates between estimating the parameters while fixing the disparity map and estimating the disparity map while fixing the parameters. The estimated parameters include robust truncation thresholds for both data and neighborhood terms, as well as a regularization weight. The regularization weight can be either a constant for the whole image or spatially-varying, depending on local intensity gradients. In the latter case, the weights for intensity gradients are also estimated. Our approach works as a wrapper for existing stereo algorithms based on graph cuts or belief propagation, automatically tuning their parameters to improve performance without requiring the stereo code to be modified. Experiments demonstrate that our approach moves a baseline belief propagation stereo algorithm up six slots in the Middlebury rankings. Li Zhang 0003, Steven M. Seitz |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Multi-View Stereo RevisitedabstractWe present an extremely simple yet robust multi-view stereo algorithm and analyze its properties. The algorithm first computes individual depth maps using a window-based voting approach that returns only good matches. The depth maps are then merged into a single mesh using a straightforward volumetric approach. We show results for several datasets, showing accuracy comparable to the best of the current state of the art techniques and rivaling more complex algorithms. Michael Goesele, Brian Curless, Steven M. Seitz |
CVPR (2) | 3 |
| 2006 | A Comparison and Evaluation of Multi-View Stereo Reconstruction AlgorithmsabstractThis paper presents a quantitative comparison of several multi-view stereo reconstruction algorithms. Until now, the lack of suitable calibrated multi-view image datasets with known ground truth (3D shape models) has prevented such direct comparisons. In this paper, we first survey multi-view stereo algorithms and compare them qualitatively using a taxonomy that differentiates their key properties. We then describe our process for acquiring and calibrating multiview image datasets with high-accuracy ground truth and introduce our evaluation methodology. Finally, we present the results of our quantitative comparison of state-of-the-art multi-view stereo reconstruction algorithms on six benchmark datasets. The datasets, evaluation details, and instructions for submitting new models are available online at http://vision.middlebury.edu/mview. Steven M. Seitz, Brian Curless, James Diebel, Daniel Scharstein, Richard Szeliski |
CVPR (1) | 1 |
| 2006 | Schematic storyboarding for video visualization and editingabstractWe present a method for visualizing short video clips in a single static image, using the visual language of storyboards. These schematic storyboards are composed from multiple input frames and annotated using outlines, arrows, and text describing the motion in the scene. The principal advantage of this storyboard representation over standard representations of video -- generally either a static thumbnail image or a playback of the video clip in its entirety -- is that it requires only a moment to observe and comprehend but at the same time retains much of the detail of the source video. Our system renders a schematic storyboard layout based on a small amount of user interaction. We also demonstrate an interaction technique to scrub through time using the natural spatial dimensions of the storyboard. Potential applications include video editing, surveillance summarization, assembly instructions, composition of graphic novels, and illustration of camera technique for film studies. Dan B. Goldman, Brian Curless, David Salesin, Steven M. Seitz |
ACM Trans. Graph. | 4 |
| 2006 | Photo tourism: exploring photo collections in 3DabstractWe present a system for interactively browsing and exploring large unstructured collections of photographs of a scene using a novel 3D interface. Our system consists of an image-based modeling front end that automatically computes the viewpoint of each photograph as well as a sparse 3D model of the scene and image to model correspondences. Our photo explorer uses image-based rendering techniques to smoothly transition between photographs, while also enabling full 3D navigation and exploration of the set of images and world geometry, along with auxiliary information such as overhead maps. Our system also makes it easy to construct photo tours of scenic or historic locations, and to annotate image details, which are automatically transferred to other relevant images. We demonstrate our system on several large personal photo collections as well as images gathered from Internet photo sharing sites. Noah Snavely, Steven M. Seitz, Richard Szeliski |
ACM Trans. Graph. | 2 |
| 2005 | Parameter Estimation for MRF StereoabstractThis paper presents a novel approach for estimating parameters for MRF-based stereo algorithms. This approach is based on a new formulation of stereo as a maximum a posterior (MAP) problem, in which both a disparity map and MRF parameters are estimated from the stereo pair itself. We present an iterative algorithm for the MAP estimation that alternates between estimating the parameters while fixing the disparity map and estimating the disparity map while fixing the parameters. The estimated parameters include robust truncation thresholds, for both data and neighborhood terms, as well as a regularization weight. The regularization weight can be either a constant for the whole image, or spatially-varying, depending on local intensity gradients. In the latter case, the weights for intensity gradients are also estimated. Experiments indicate that our approach, as a wrapper for existing stereo algorithms, moves a baseline belief propagation stereo algorithm up six slots in the middlebury rankings. Li Zhang 0003, Steven M. Seitz |
CVPR (2) | 2 |
| 2005 | Shape and Spatially-Varying BRDFs from Photometric StereoabstractThis paper describes a photometric stereo method designed for surfaces with spatially-varying BRDFs, including surfaces with both varying diffuse and specular properties. Our method builds on the observation that most objects are composed of a small number of fundamental materials. This approach recovers not only the shape but also material BRDFs and weight maps, yielding compelling results for a wide variety of objects. We also show examples of interactive lighting and editing operations made possible by our method. Dan B. Goldman, Brian Curless, Aaron Hertzmann, Steven M. Seitz |
ICCV | 4 |
| 2005 | A Theory of Inverse Light TransportabstractIn this paper we consider the problem of computing and removing interreflections in photographs of real scenes. Towards this end, we introduce the problem of inverse light transport - given a photograph of an unknown scene, decompose it into a sum of n-bounce images, where each image records the contribution of light that bounces exactly n times before reaching the camera. We prove the existence of a set of interreflection cancelation operators that enable computing each n-bounce image by multiplying the photograph by a matrix. This matrix is derived from a set of "impulse images" obtained by probing the scene with a narrow beam of light. The operators work under unknown and arbitrary illumination, and exist for scenes that have arbitrary spatially-varying BRDFs. We derive a closed-form expression for these operators in the Lambertian case and present experiments with textured and untextured Lambertian scenes that confirm our theory's predictions. Steven M. Seitz, Yasuyuki Matsushita, Kiriakos N. Kutulakos |
ICCV | 1 |
| 2005 | Example-Based Photometric Stereo: Shape Reconstruction with General, Varying BRDFsabstractThis paper presents a technique for computing the geometry of objects with general reflectance properties from images. For surfaces with varying material properties, a full segmentation into different material types is also computed. It is assumed that the camera viewpoint is fixed, but the illumination varies over the input sequence. It is also assumed that one or more example objects with similar materials and known geometry are imaged under the same illumination conditions. Unlike most previous work in shape reconstruction, this technique can handle objects with arbitrary and spatially-varying BRDFs. Furthermore, the approach works for arbitrary distant and unknown lighting environments. Finally, almost no calibration is needed, making the approach exceptionally simple to apply. Aaron Hertzmann, Steven M. Seitz |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Reconstructing Relief SurfacesabstractThis paper generalizes Markov Random Field (MRF) stereo methods to the generation of surface relief (height) fields rather than disparity or depth maps. This generalization enables the reconstruction of complete object models using the same algorithms that have been previously used to compute depth maps in binocular stereo. In contrast to traditional dense stereo where the parametrization is image based, here we advocate a parametrization by a height field over any base surface. In practice, the base surface is a coarse approximation to the true geometry, e.g., a bounding box, visual hull or triangulation of sparse correspondences, and is assigned or computed using other means. A dense set of sample points is defined on the base surface, each with a fixed normal direction and unknown height value. The estimation of heights for the sample points is achieved by a belief propagation technique. Our method provides a viewpoint independent smoothness constraint, a more compact parametrization and explicit handling of occlusions. We present experimental results on real scenes as well as a quantitative evaluation on an artificial scene. George Vogiatzis, Philip Torr 0001, Steven M. Seitz, Roberto Cipolla |
BMVC | 3 |
| 2004 | Example-Based Stereo with General BRDFs
Adrien Treuille, Aaron Hertzmann, Steven M. Seitz |
ECCV (2) | 3 |
| 2004 | Video-based document tracking: unifying your physical and electronic desktopsabstractThis paper presents an approach for tracking paper documents on the desk over time and automatically linking them to the corresponding electronic documents using an overhead video camera. We demonstrate our system in the context of two scenarios, paper tracking and photo sorting. In the paper tracking scenario, the system tracks changes in the stacks of printed documents and books on the desk and builds a complete representation of the spatial structure of the desktop. When users want to find a printed document buried in the stacks, they can query the system based on appearance, keywords, or access time. The system also provides a remote desktop interface for directly browsing the physical desktop from a remote location. In the photo sorting scenario, users sort printed photographs into physical stacks on the desk. The systemautomatically recognizes the photographs and organizes the corresponding digital photographs into separate folders according to the physical arrangement. Our framework provides a way to unify the physical and electronic desktops without the need for a specialized physical infrastructure except for a video camera. Steven M. Seitz, Maneesh Agrawala |
UIST | 2 |
| 2004 | Keyframe-based tracking for rotoscoping and animationabstractWe describe a new approach to rotoscoping --- the process of tracking contours in a video sequence --- that combines computer vision with user interaction. In order to track contours in video, the user specifies curves in two or more frames; these curves are used as keyframes by a computer-vision-based tracking algorithm. The user may interactively refine the curves and then restart the tracking algorithm. Combining computer vision with user interaction allows our system to track any sequence with significantly less effort than interpolation-based systems --- and with better reliability than "pure" computer vision systems. Our tracking algorithm is cast as a spacetime optimization problem that solves for time-varying curve shapes based on an input video sequence and user-specified constraints. We demonstrate our system with several rotoscoped examples. Additionally, we show how these rotoscoped contours can be used to help create cartoon animation by attaching user-drawn strokes to the tracked contours. Aseem Agarwala, Aaron Hertzmann, David Salesin, Steven M. Seitz |
ACM Trans. Graph. | 4 |
| 2004 | Flow-based video synthesis and editingabstractThis paper presents a novel algorithm for synthesizing and editing video of natural phenomena that exhibit continuous flow patterns. The algorithm analyzes the motion of textured particles in the input video along user-specified flow lines, and synthesizes seamless video of arbitrary length by enforcing temporal continuity along a second set of user-specified flow lines. The algorithm is simple to implement and use. We used this technique to edit video of water-falls, rivers, flames, and smoke. Kiran S. Bhat, Steven M. Seitz, Jessica K. Hodgins, Pradeep K. Khosla |
ACM Trans. Graph. | 2 |
| 2004 | Spacetime faces: high resolution capture for modeling and animationabstractWe present an end-to-end system that goes from video sequences to high resolution, editable, dynamically controllable face models. The capture system employs synchronized video cameras and structured light projectors to record videos of a moving face from multiple viewpoints. A novel spacetime stereo algorithm is introduced to compute depth maps accurately and overcome over-fitting deficiencies in prior work. A new template fitting and tracking procedure fills in missing data and yields point correspondence across the entire sequence without using markers. We demonstrate a data-driven, interactive method for inverse kinematics that draws on the large set of fitted templates and allows for posing new expressions by dragging surface points directly. Finally, we describe new tools that model the dynamics in the input sequence to enable new animations, created via key-framing or texture-synthesis techniques. Li Zhang 0003, Noah Snavely, Brian Curless, Steven M. Seitz |
ACM Trans. Graph. | 4 |
| 2003 | Shape and Materials by Example: A Photometric Stereo ApproachabstractThis paper presents a technique for computing the geometry of objects with general reflectance properties from images. For surfaces with varying material properties, a full segmentation into different material types is also computed. It is assumed that the camera viewpoint is fixed, but the illumination varies over the input sequence. It is also assumed that one or more example objects with similar materials and known geometry are imaged under the same illumination conditions. Unlike most previous work in shape reconstruction, this technique can handle objects with arbitrary and spatially-varying BRDFs. Furthermore, the approach works for arbitrary distant and unknown lighting environments. Finally, almost no calibration is needed, making the approach exceptionally simple to apply. Aaron Hertzmann, Steven M. Seitz |
CVPR (1) | 2 |
| 2003 | Spacetime Stereo: Shape Recovery for Dynamic ScenesabstractThis paper extends the traditional binocular stereo problem into the spacetime domain, in which a pair of video streams is matched simultaneously instead of matching pairs of images frame by frame. Almost any existing stereo algorithm may be extended in this manner simply by replacing the image matching term with a spacetime term. By utilizing both spatial and temporal appearance variation, this modification reduces ambiguity and increases accuracy. Three major applications for spacetime stereo are proposed in this paper. First, spacetime stereo serves as a general framework for structured light scanning and generates high quality depth maps for static scenes. Second, spacetime stereo is effective for a class of natural scenes, such as waving trees and flowing water, which have repetitive textures and chaotic behaviors and are challenging for existing stereo algorithms. Third, the approach is one of very few existing methods that can robustly reconstruct objects that are moving and deforming over time, achieved by use of oriented spacetime windows in the matching procedure. Promising experimental results in the above three scenarios are demonstrated. Li Zhang 0003, Brian Curless, Steven M. Seitz |
CVPR (2) | 3 |
| 2003 | Shape and Motion under Varying Illumination: Unifying Structure from Motion, Photometric Stereo, and Multi-view StereoabstractWe present an algorithm for computing optical flow, shape, motion, lighting, and albedo from an image sequence of a rigidly-moving Lambertian object under distant illumination. The problem is formulated in a manner that subsumes structure from motion, multiview stereo, and photometric stereo as special cases. The algorithm utilizes both spatial and temporal intensity variation as cues: the former constrains flow and the latter constrains surface orientation; combining both cues enables dense reconstruction of both textured and textureless surfaces. The algorithm works by iteratively estimating affine camera parameters, illumination, shape, and albedo in an alternating fashion. Results are demonstrated on videos of hand-held objects moving in front of a fixed light and camera. Li Zhang 0003, Brian Curless, Aaron Hertzmann, Steven M. Seitz |
ICCV | 4 |
| 2003 | EM, MCMC, and Chain Flipping for Structure from Motion with Unknown Correspondence
Frank Dellaert, Steven M. Seitz, Charles E. Thorpe, Sebastian Thrun |
Mach. Learn. | 2 |
| 2003 | Motion sketching for control of rigid-body simulationsabstractMotion sketching is an approach for creating realistic rigid-body motion. In this approach, an animator sketches how objects should move and the system computes a physically plausible motion that best fits the sketch. The sketch is specified with a mouse-based interface or with hand-gestures, which move instrumented objects in the real world to act out the desired behaviors. The sketches may be imprecise, may be physically infeasible, or may have incorrect timing. A multiple-shooting optimization estimates the parameters of a rigid-body simulation needed to simulate an animation that matches the sketch with physically plausible timing and motion. This technique applies to physical simulations of multiple colliding rigid bodies possibly connected with joints in a tree (open-loop) topology. Jovan Popovic, Steven M. Seitz, Michael A. Erdmann |
ACM Trans. Graph. | 2 |
| 2002 | Computing the Physical Parameters of Rigid-Body Motion from Video
Kiran S. Bhat, Steven M. Seitz, Jovan Popovic |
ECCV (1) | 2 |
| 2002 | Techniques for Interactive Audience ParticipationabstractAt SIGGRAPH in 1991, Loren and Rachel Carpenter unveiled an interactive entertainment system that allowed members of a large audience to control an onscreen game using red and green reflective paddles. In the spirit of this approach, we present a new set of techniques that enable members of an audience to participate, either cooperatively or competitively, in shared entertainment experiences. Our techniques allow audiences with hundreds of people to control onscreen activity by (1) leaning left and right in their seats, (2) batting a beach ball while its shadow is used as a pointing device, and (3) pointing laser pointers at the screen. All of these techniques can be implemented with inexpensive, off the shelf hardware. Me have tested these techniques with a variety of audiences; in this paper we describe both the computer vision based implementation and the lessons we learned about designing effective content for interactive audience participation. Dan Maynes-Aminzade, Randy F. Pausch, Steven M. Seitz |
ICMI | 3 |
| 2002 | The Space of All Stereo Images
Steven M. Seitz |
Int. J. Comput. Vis. | 1 |
| 2002 | Plenoptic Image Editing
Steven M. Seitz, Kiriakos N. Kutulakos |
Int. J. Comput. Vis. | 1 |
| 2002 | Omnivergent Stereo
Steven M. Seitz, Adam Tauman Kalai, Harry Shum |
Int. J. Comput. Vis. | 1 |
| 2002 | Single-view modelling of free-form scenesabstractAbstract This paper presents a novel approach for reconstructing free‐form, texture‐mapped, 3D scene models from a single painting or photograph. Given a sparse set of user‐specified constraints on the local shape of the scene, a smooth 3D surface that satisfies the constraints is generated. This problem is formulated as a constrained variational optimization problem. In contrast to previous work in single‐view reconstruction, our technique enables high‐quality reconstructions of free‐form curved surfaces with arbitrary reflectance properties. A key feature of the approach is a novel hierarchical transformation technique for accelerating convergence on a non‐uniform, piecewise continuous grid. The technique is interactive and updates the model in real time as constraints are added, allowing fast reconstruction of photorealistic scene models. The approach is shown to yield high‐quality results on a large variety of images. Copyright © 2002 John Wiley & Sons, Ltd. Li Zhang 0003, Guillaume Dugas-Phocion, Jean-Sebastien Samson, Steven M. Seitz |
Comput. Animat. Virtual Worlds | 4 |
| 2001 | Single View Modeling of Free-Form ScenesabstractThis paper presents a novel approach for reconstructing free-form, texture-mapped, 3D scene models from a single painting or photograph. Given a sparse set of user-specified constraints on the local shape of the scene, a smooth 3D surface that satisfies the constraints is generated This problem is formulated as a constrained variational optimization problem. In contrast to previous work in single view reconstruction, our technique enables high quality reconstructions of free-form curved surfaces with arbitrary reflectance properties. A key feature of the approach is a novel hierarchical transformation technique for accelerating convergence on a non-uniform, piecewise continuous grid. The technique is interactive and updates the model in real time as constraints are added, allowing fast reconstruction of photorealistic scene models. The approach is shown to yield high quality results on a large variety of images. Li Zhang 0003, Guillaume Dugas-Phocion, Jean-Sebastien Samson, Steven M. Seitz |
CVPR (1) | 4 |
| 2001 | The Space of All Stereo Images
Steven M. Seitz |
ICCV | 1 |
| 2000 | Structure from Motion without CorrespondenceabstractA method is presented to recover 3D scene structure and camera motion from multiple images without the need for correspondence information. The problem is framed as finding the maximum likelihood structure and motion given only the 2D measurements, integrating over all possible assignments of 3D features to 2D measurements. This goal is achieved by means of an algorithm which iteratively refines a probability distribution over the set of all correspondence assignments. At each iteration a new structure from motion problem is solved, using as input a set of 'virtual measurements' derived from this probability distribution. The distribution needed can be efficiently obtained by Markov Chain Monte Carlo sampling. The approach is cast within the framework of Expectation-Maximization, which guarantees convergence to a local maximizer of the likelihood. The algorithm works well in practice, as will be demonstrated using results on several real image sequences. Frank Dellaert, Steven M. Seitz, Charles E. Thorpe, Sebastian Thrun |
CVPR | 2 |
| 2000 | Visual Tunnel Analysis for Visibility Prediction and Camera PlanningabstractA sequence of images taken along a camera trajectory captures a subset of scene appearance. If visibility space is the space that encapsulates the appearance of the scene at every conceivable pose and viewing angle, then the act of acquiring the image sequence constitutes "carving a volume in visibility space." We call such a volume a visual tunnel. The analysis of the visual tunnel allows us to do the following: predict the range of virtual camera poses in which the images can be reconstructed totally using the captured rays, predict which parts of the image can be generated for a given virtual camera pose, and plan camera paths for scene visualization at desired locations. We describe our visual tunnel concept and provide illustrative examples in 2D and 3D. Sing Bing Kang, Peter-Pike J. Sloan, Steven M. Seitz |
CVPR | 3 |
| 2000 | Shape and Motion Carving in 6DabstractThe motion of a non-rigid scene over time imposes more constraints on its structure than those derived from images at a single time instant alone. An algorithm is presented for simultaneously recovering dense scene shape and scene flow (i.e. the instantaneous 3D motion at every point in the scene). The algorithm operates by carving away hexels, or points in the 6D space of all possible shapes and flows that are inconsistent with the images captures at either time instant, or across time. The recovered shape is demonstrated to be more accurate than that recovered using images at a single time instant. Applications of the combined scene shape and flow include motion capture for animation, retiming of videos, and non-rigid motion analysis. Sundar Vedula, Simon Baker, Steven M. Seitz, Takeo Kanade |
CVPR | 3 |
| 2000 | Feature Correspondence: A Markov Chain Monte Carlo ApproachabstractWhen trying to recover 3D structure from a set of images, the most difficult problem is establishing the correspondence between the measurements. Most existing approaches assume that features can be tracked across frames, whereas methods that exploit rigidity constraints to facilitate matching do so only under restricted cam(cid:173) era motion. In this paper we propose a Bayesian approach that avoids the brittleness associated with singling out one "best" cor(cid:173) respondence, and instead consider the distribution over all possible correspondences. We treat both a fully Bayesian approach that yields a posterior distribution, and a MAP approach that makes use of EM to maximize this posterior. We show how Markov chain Monte Carlo methods can be used to implement these techniques in practice, and present experimental results on real data. Frank Dellaert, Steven M. Seitz, Sebastian Thrun, Charles E. Thorpe |
NIPS | 2 |
| 2000 | Interactive manipulation of rigid body simulationsabstractPhysical simulation of dynamic objects has become commonplace in computer graphics because it produces highly realistic animations. In this paradigm the animator provides few physical parameters such as the objects' initial positions and velocities, and the simulator automatically generates realistic motions. The resulting motion, however, is difficult to control because even a small adjustment of the input parameters can drastically affect the subsequent motion. Furthermore, the animator often wishes to change the end-result of the motion instead of the initial physical parameters. Jovan Popovic, Steven M. Seitz, Michael A. Erdmann, Zoran Popovic, Andrew P. Witkin |
SIGGRAPH | 2 |
| 2000 | A Theory of Shape by Space Carving
Kiriakos N. Kutulakos, Steven M. Seitz |
Int. J. Comput. Vis. | 2 |
| 1999 | Implicit Representation and Scene Reconstruction from Probability Density FunctionsabstractA technique is presented for representing linear features as probability density functions in two or three dimensions. Three chief advantages of this approach are (1) a unified representation and algebra for manipulating points, lines, and planes, (2) seamless incorporation of uncertainty information, and (3) a very simple recursive solution for maximum likelihood shape estimation. Applications to uncalibrated affine scene reconstruction are presented, with results on images of an outdoor environment. Steven M. Seitz, P. Anandan 0001 |
CVPR | 1 |
| 1999 | A Theory of Shape by Space CarvingabstractIn this paper we consider the problem of computing the 3D shape of an unknown, arbitrarily-shaped scene from multiple photographs taken at known but arbitrarily-distributed viewpoints. By studying the equivalence class of all 3D shapes that reproduce the input photographs, we prove the existence of a special member of this class, the photo hull, that (1) can be computed directly from photographs of the scene, and (2) subsumes all other members of this class. We then give a provably-correct algorithm called Space Carving, for computing this shape and present experimental results on complex real-world scenes. The approach is designed to (1) build photorealistic shapes that accurately model scene appearance from a wide range of viewpoints, and (2) account for the complex interactions between occlusion, parallax, shading, and their effects on arbitrary views of a 3D scene. Kiriakos N. Kutulakos, Steven M. Seitz |
ICCV | 2 |
| 1999 | Omnivergent StereoabstractThe notion of a virtual sensor for optimal 3D reconstruction is introduced. Instead of planar perspective images that collect many rays at a fixed viewpoint, omnivergent cameras collect a small number of rays at many different viewpoints. The resulting 2D manifold of rays are arranged into two multiple-perspective images for stereo reconstruction. We call such images omnivergent images, and the process of reconstructing the scene from such images, omnivergent stereo. This procedure is shown to produce 3D scene models with minimal reconstruction error due to the fact that for any point in the 3D scene, two rays with maximum vergence angle can be found in the omnivergent images. Furthermore, omnivergent images are shown to have horizontal epipolar lines, enabling the application of traditional stereo matching algorithms, without modification. Three types of omnivergent virtual sensors are presented: spherical omnivergent cameras, center-strip cameras and dual-strip cameras. Harry Shum, Adam Tauman Kalai, Steven M. Seitz |
ICCV | 3 |
| 1999 | Photorealistic Scene Reconstruction by Voxel Coloring
Steven M. Seitz, Charles R. Dyer |
Int. J. Comput. Vis. | 1 |
| 1998 | Plenoptic Image EditingabstractThis paper presents a new class of interactive image editing operations designed to maintain consistency between multiple images of a physical 3D scene. The distinguishing feature of these operations is that edits to any one image propagate automatically to all other images as if the (unknown) 3D scene had itself been modified. The modified scene can then be viewed interactively from any other camera viewpoint and under different scene illuminations. The approach is useful first as a power-assist that enables a user to quickly modify many images by editing just a few, and second as a means for constructing and editing image-based scene representations by manipulating a set of photographs. The approach works by extending operations like image painting, scissoring, and morphing so that they alter a scene's generalized plenoptic function in a physically-consistent way, thereby affecting scene appearance from all viewpoints simultaneously. A key element in realizing these operations is a new volumetric decomposition technique for reconstructing an scene's plenoptic function from an incomplete set of camera viewpoints. Steven M. Seitz, Kiriakos N. Kutulakos |
ICCV | 1 |
| 1997 | Photorealistic Scene Reconstruction by Voxel ColoringabstractA novel scene reconstruction technique is presented, different from previous approaches in its ability to cope with large changes in visibility and its modeling of intrinsic scene color and texture information. The method avoids image correspondence problems by working in a discretized scene space whose voxels are traversed in a fixed visibility ordering. This strategy takes full account of occlusions and allows the input cameras to be far apart and widely distributed about the environment. The algorithm identifies a special set of invariant voxels which together form a spatial and photometric reconstruction of the scene, fully consistent with the input images. The approach is evaluated with images from both inward- and outward-facing cameras. Steven M. Seitz, Charles R. Dyer |
CVPR | 1 |
| 1997 | View-Invariant Analysis of Cyclic Motion
Steven M. Seitz, Charles R. Dyer |
Int. J. Comput. Vis. | 1 |
| 1996 | Toward image-based scene representation using view morphingabstractThe question of which views map be inferred from a set of basis images is addressed. Under certain conditions, a discrete set of images implicitly describes scene appearance for a continuous range of viewpoints. In particular it is demonstrated that two basis views of a static scene determine the set of all views an the line between their optical centers. Additional basis views further extend the range of predictable views to a two- or three-dimensional region of viewspace. These results are shown to apply under perspective projection subject to a generic visibility constraint called monotonicity. In addition, a simple scanline algorithm is presented for actually generating these views from a set of basis images. The technique, called view morphing map be applied to both calibrated and uncalibrated images. At a minimum, two basis views and their fundamental matrix are needed. Experimental results are presented an real images. This work provides a theoretical foundation for image-based representations of 3D scenes by demonstrating that perspective view synthesis is a theoretically well-posed problem. Steven M. Seitz, Charles R. Dyer |
ICPR | 1 |
| 1996 | View MorphingabstractArticle View morphing Share on Authors: Steven M. Seitz Department of Computer Sciences, University of Wisconsin--Madison, 1210 W. Dayton St., Madison WI Department of Computer Sciences, University of Wisconsin--Madison, 1210 W. Dayton St., Madison WIView Profile , Charles R. Dyer Department of Computer Sciences, University of Wisconsin--Madison, 1210 W. Dayton St., Madison WI Department of Computer Sciences, University of Wisconsin--Madison, 1210 W. Dayton St., Madison WIView Profile Authors Info & Claims SIGGRAPH '96: Proceedings of the 23rd annual conference on Computer graphics and interactive techniquesAugust 1996 Pages 21–30https://doi.org/10.1145/237170.237196Online:01 August 1996Publication History 471citation3,015DownloadsMetricsTotal Citations471Total Downloads3,015Last 12 Months115Last 6 weeks4 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access Steven M. Seitz, Charles R. Dyer |
SIGGRAPH | 1 |
| 1995 | Complete Scene Structure from Four Point CorrespondencesabstractA technique is presented for computing 3D scene structure from point and line features in monocular image sequences. Unlike previous methods, the technique guarantees the completeness of the recovered scene, ensuring that every scene feature that is detected in each image is reconstructed. The approach relies on the presence of four or more reference features whose correspondences are known in all the images. Under an orthographic or affine camera model, the parallax of the reference features provides constraints that simplify the recovery of the rest of the visible scene. An efficient recursive algorithm is described that uses a unified framework for point and line features. The algorithm integrates the tasks of feature correspondence and structure recovery, ensuring that all reconstructible features are tracked. In addition, the algorithm is immune to outliers and feature drift, two weaknesses of existing structure from motion techniques. Experimental results are presented for real images.> Steven M. Seitz, Charles R. Dyer |
ICCV | 1 |
| 1994 | Affine invariant detection of periodic motionabstractCurrent approaches for detecting periodic motion assume a stationary camera and place limits on an object's motion. These approaches rely on the assumption that a periodic motion projects to a set of periodic image curves, an assumption that fails in general. Using affine-invariance, we derive necessary and sufficient conditions for an image sequence to be the projection of a periodic motion. No restrictions are placed on either the motion of the camera or the object. Our algorithm is shown to be provably-correct for noise-free data and is easily extended to be robust with respect to occlusions and noise. The extended algorithm is evaluated with real and synthetic image sequences.> Steven M. Seitz, Charles R. Dyer |
CVPR | 1 |