Richard Szeliski

dblp:46/2186 · DBLP profile ↗
← Back
165ranked-venue papers
38as first author
8since 2021 · last 2024
0009-0005-5300-5475ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 124 · 27 first-author · 7 since 2021Artificial intelligence and machine learning · 111 · 29 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 11 · 3 first-author · 2 since 2021Computer networks · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2024 UniSDF: Unifying Neural Representations for High-Fidelity 3D Reconstruction of Complex Scenes with Reflections
abstract
Neural 3D scene representations have shown great potential for 3D reconstruction from 2D images. However, reconstructing real-world captures of complex scenes still remains a challenge. Existing generic 3D reconstruction methods often struggle to represent fine geometric details and do not adequately model reflective surfaces of large-scale scenes. Techniques that explicitly focus on reflective surfaces can model complex and detailed reflections by exploiting better reflection parameterizations. However, we observe that these methods are often not robust in real scenarios where non-reflective as well as reflective components are present. In this work, we propose UniSDF, a general purpose 3D reconstruction method that can reconstruct large complex scenes with reflections. We investigate both camera view as well as reflected view-based color parameterization techniques and find that explicitly blending these representations in 3D space enables reconstruction of surfaces that are more geometrically accurate, especially for reflective surfaces. We further combine this representation with a multi-resolution grid backbone that is trained in a coarse-to-fine manner, enabling faster reconstructions than prior methods. Extensive experiments on object-level datasets DTU, Shiny Blender as well as unbounded datasets Mip-NeRF 360 and Ref-NeRF real demonstrate that our method is able to robustly reconstruct complex large-scale scenes with fine details and reflective surfaces, leading to the best overall performance. Project page: https://fangjinhuawang.github.io/UniSDF.
Fangjinhua Wang, Marie-Julie Rakotosaona, Michael Niemeyer, Richard Szeliski, Marc Pollefeys, Federico Tombari
NeurIPS4
2024 A Plentoptic 3D Vision System
Agastya Kalra, Vage Taamazyan, Alberto Dall'olio, Raghav Khanna, Tomas Gerlich, Georgia Giannopolou, Guy Stoppi, Daniel Baxter, Abhijit Ghosh, Richard Szeliski, Kartik Venkataraman
SIGGRAPH Asia10
2024 NeRF-Casting: Improved View-Dependent Appearance with Consistent Reflections
Dor Verbin, Pratul P. Srinivasan, Peter Hedman, Ben Mildenhall, Benjamin Attal, Richard Szeliski, Jonathan T. Barron
SIGGRAPH Asia6
2024 SMERF: Streamable Memory Efficient Radiance Fields for Real-Time Large-Scene Exploration
abstract
Recent techniques for real-time view synthesis have rapidly advanced in fidelity and speed, and modern methods are capable of rendering near-photorealistic scenes at interactive frame rates. At the same time, a tension has arisen between explicit scene representations amenable to rasterization and neural fields built on ray marching, with state-of-the-art instances of the latter surpassing the former in quality while being prohibitively expensive for real-time applications. We introduce SMERF, a view synthesis approach that achieves state-of-the-art accuracy among real-time methods on large scenes with footprints up to 300 m 2 at a volumetric resolution of 3.5 mm 3 . Our method is built upon two primary contributions: a hierarchical model partitioning scheme, which increases model capacity while constraining compute and memory consumption, and a distillation training strategy that simultaneously yields high fidelity and internal consistency. Our method enables full six degrees of freedom navigation in a web browser and renders in real-time on commodity smartphones and laptops. Extensive experiments show that our method exceeds the state-of-the-art in real-time novel view synthesis by 0.78 dB on standard benchmarks and 1.78 dB on large scenes, renders frames three orders of magnitude faster than state-of-the-art radiance field models, and achieves real-time performance across a wide variety of commodity devices, including smartphones. We encourage readers to explore these models interactively at our project website: https://smerf-3d.github.io.
Daniel Duckworth, Peter Hedman, Christian Reiser, Peter Zhizhin, Jean-François Thibert, Mario Lucic, Richard Szeliski, Jonathan T. Barron
ACM Trans. Graph.7
2024 Binary Opacity Grids: Capturing Fine Geometric Detail for Mesh-Based View Synthesis
abstract
While surface-based view synthesis algorithms are appealing due to their low computational requirements, they often struggle to reproduce thin structures. In contrast, more expensive methods that model the scene's geometry as a volumetric density field (e.g. NeRF) excel at reconstructing fine geometric detail. However, density fields often represent geometry in a "fuzzy" manner, which hinders exact localization of the surface. In this work, we modify density fields to encourage them to converge towards surfaces, without compromising their ability to reconstruct thin structures. First, we employ a discrete opacity grid representation instead of a continuous density field, which allows opacity values to discontinuously transition from zero to one at the surface. Second, we anti-alias by casting multiple rays per pixel, which allows occlusion boundaries and subpixel structures to be modelled without using semi-transparent voxels. Third, we minimize the binary entropy of the opacity values, which facilitates the extraction of surface geometry by encouraging opacity values to binarize towards the end of training. Lastly, we develop a fusion-based meshing strategy followed by mesh simplification and appearance model fitting. The compact meshes produced by our model can be rendered in real-time on mobile devices and achieve significantly higher view synthesis quality compared to existing mesh-based approaches. Our interactive webdemo is available at https://binary-opacity-grid.github.io.
Christian Reiser, Stephan J. Garbin, Pratul P. Srinivasan, Dor Verbin, Richard Szeliski, Ben Mildenhall, Jonathan T. Barron, Peter Hedman, Andreas Geiger 0001
ACM Trans. Graph.5
2023 Accidental Light Probes
abstract
Recovering lighting in a scene from a single image is a fundamental problem in computer vision. While a mirror ball light probe can capture omnidirectional lighting, light probes are generally unavailable in everyday images. In this work, we study recovering lighting from accidental light probes (ALPs)-common, shiny objects like Coke cans, which often accidentally appear in daily scenes. We propose a physically-based approach to model ALPs and estimate lighting from their appearances in single images. The main idea is to model the appearance of ALPs by photogram-metrically principled shading and to invert this process via differentiable rendering to recover incidental illumination. We demonstrate that we can put an ALP into a scene to allow high-fidelity lighting estimation. Our model can also recover lighting for existing images that happen to contain an ALP**Project website: https://kovenyu.com/ALP. I'd rather be Shiny. - Tamatoa from Moana, 2016
Hong-Xing Yu, Samir Agarwala, Charles Herrmann, Richard Szeliski, Noah Snavely, Jiajun Wu 0001, Deqing Sun
CVPR4
2023 MERF: Memory-Efficient Radiance Fields for Real-time View Synthesis in Unbounded Scenes
abstract
Neural radiance fields enable state-of-the-art photorealistic view synthesis. However, existing radiance field representations are either too compute-intensive for real-time rendering or require too much memory to scale to large scenes. We present a Memory-Efficient Radiance Field (MERF) representation that achieves real-time rendering of large-scale scenes in a browser. MERF reduces the memory consumption of prior sparse volumetric radiance fields using a combination of a sparse feature grid and high-resolution 2D feature planes. To support large-scale unbounded scenes, we introduce a novel contraction function that maps scene coordinates into a bounded volume while still allowing for efficient ray-box intersection. We design a lossless procedure for baking the parameterization used during training into a model that achieves real-time rendering while still preserving the photorealistic view synthesis quality of a volumetric radiance field.
Christian Reiser, Richard Szeliski, Dor Verbin, Pratul P. Srinivasan, Ben Mildenhall, Andreas Geiger 0001, Jonathan T. Barron, Peter Hedman
ACM Trans. Graph.2
2021 Animating Pictures With Eulerian Motion Fields
abstract
In this paper, we demonstrate a fully automatic method for converting a still image into a realistic animated looping video. We target scenes with continuous fluid motion, such as flowing water and billowing smoke. Our method relies on the observation that this type of natural motion can be convincingly reproduced from a static Eulerian motion description, i.e. a single, temporally constant flow field that defines the immediate motion of a particle at a given 2D location. We use an image-to-image translation network to encode motion priors of natural scenes collected from on-line videos, so that for a new photo, we can synthesize a corresponding motion field. The image is then animated using the generated motion through a deep warping technique: pixels are encoded as deep features, those features are warped via Eulerian motion, and the resulting warped feature maps are decoded as images. In order to produce continuous, seamlessly looping video textures, we propose a novel video looping technique that flows features both for-ward and backward in time and then blends the results. We demonstrate the effectiveness and robustness of our method by applying it to a large collection of examples including beaches, waterfalls, and flowing rivers.
Aleksander Holynski, Brian Curless, Steven M. Seitz, Richard Szeliski
CVPR4
2020 Reducing Drift in Structure From Motion Using Extended Features
abstract
Low-frequency long-range errors (drift) are an endemic problem in 3D structure from motion, and can often hamper reasonable reconstructions of the scene. In this paper, we present a method to dramatically reduce scale and positional drift by using extended structural features such as planes and vanishing points. Unlike traditional feature matches, our extended features are able to span non-overlapping input images, and hence provide long-range constraints on the scale and shape of the reconstruction. We add these features as additional constraints to a state-of the-art global structure from motion algorithm and demonstrate that the added constraints enable the reconstruction of particularly drift-prone sequences such as long, low field-of-view videos without inertial measurements. Additionally, we provide an analysis of the drift-reducing capabilities of these constraints by evaluating on a synthetic dataset. Our structural features are able to significantly reduce drift for scenes that contain long-spanning man-made structures, such as aligned rows of windows or planar building facades.
Aleksander Holynski, David Geraghty, Jan-Michael Frahm, Chris Sweeney, Richard Szeliski
3DV5
2020 VPLNet: Deep Single View Normal Estimation With Vanishing Points and Lines
abstract
We present a novel single-view surface normal estimation method that combines traditional line and vanishing point analysis with a deep learning approach. Starting from a color image and a Manhattan line map, we use a deep neural network to regress on a dense normal map, and a dense Manhattan label map that identifies planar regions aligned with the Manhattan directions. We fuse the normal map and label map in a fully differentiable manner to produce a refined normal map as final output. To do so, we softly decompose the output into a Manhattan part and a non-Manhattan part. The Manhattan part is treated by discrete classification and vanishing points, while the non-Manhattan part is learned by direct supervision. Our method achieves state-of-the-art results on standard single-view normal estimation benchmarks. More importantly, we show that by using vanishing points and lines, our method has better generalization ability than existing works. In addition, we demonstrate how our surface normal network can improve the performance of depth estimation networks, both quantitatively and qualitatively, in particular, in 3D reconstructions of walls and other flat surfaces.
Rui Wang 0071, David Geraghty, Kevin Matzen, Richard Szeliski, Jan-Michael Frahm
CVPR4
2020 SynSin: End-to-End View Synthesis From a Single Image
abstract
View synthesis allows for the generation of new views of a scene given one or more images. This is challenging; it requires comprehensively understanding the 3D scene from images. As a result, current methods typically use multiple images, train on ground-truth depth, or are limited to synthetic data. We propose a novel end-to-end model for this task using a single image at test time; it is trained on real images without any ground-truth 3D information. To this end, we introduce a novel differentiable point cloud renderer that is used to transform a latent 3D point cloud of features into the target view. The projected features are decoded by our refinement network to inpaint missing regions and generate a realistic output image. The 3D component inside of our generative model allows for interpretable manipulation of the latent feature space at test time, e.g. we can animate trajectories from a single image. Additionally, we can generate high resolution images and generalise to other input resolutions. We outperform baselines and prior work on the Matterport, Replica, and RealEstate10K datasets.
Olivia Wiles, Georgia Gkioxari, Richard Szeliski, Justin Johnson 0001
CVPR3
2020 Consistent video depth estimation
abstract
We present an algorithm for reconstructing dense, geometrically consistent depth for all pixels in a monocular video. We leverage a conventional structure-from-motion reconstruction to establish geometric constraints on pixels in the video. Unlike the ad-hoc priors in classical reconstruction, we use a learning-based prior, i.e., a convolutional neural network trained for single-image depth estimation. At test time, we fine-tune this network to satisfy the geometric constraints of a particular input video, while retaining its ability to synthesize plausible depth details in parts of the video that are less constrained. We show through quantitative validation that our method achieves higher accuracy and a higher degree of geometric consistency than previous monocular reconstruction methods. Visually, our results appear more stable. Our algorithm is able to handle challenging hand-held captured input videos with a moderate degree of dynamic motion. The improved quality of the reconstruction enables several applications, such as scene reconstruction and advanced video-based visual effects.
Jia-Bin Huang 0001, Richard Szeliski, Kevin Matzen, Johannes Kopf 0001
ACM Trans. Graph.3
2019 Multi-frame stereo matching with edges, planes, and superpixels
Tianfan Xue, Andrew Owens, Daniel Scharstein, Michael Goesele, Richard Szeliski
Image Vis. Comput.5
2019 An integrated 6DoF video camera and system design
abstract
Designing a fully integrated 360° video camera supporting 6DoF head motion parallax requires overcoming many technical hurdles, including camera placement, optical design, sensor resolution, system calibration, real-time video capture, depth reconstruction, and real-time novel view synthesis. While there is a large body of work describing various system components, such as multi-view depth estimation, our paper is the first to describe a complete, reproducible system that considers the challenges arising when designing, building, and deploying a full end-to-end 6DoF video camera and playback environment. Our system includes a computational imaging software pipeline supporting online markerless calibration, high-quality reconstruction, and real-time streaming and rendering. Most of our exposition is based on a professional 16-camera configuration, which will be commercially available to film producers. However, our software pipeline is generic and can handle a variety of camera geometries and configurations. The entire calibration and reconstruction software pipeline along with example datasets is open sourced to encourage follow-up research in high-quality 6DoF video reconstruction and rendering 1 .
Albert Parra Pozo, Michael Toksvig, Terry Filiba Schrager, Joyce Hsu, Uday Mathur, Alexander Sorkine-Hornung, Richard Szeliski, Brian Cabral
ACM Trans. Graph.7
2018 Reconstructing scenes with mirror and glass surfaces
abstract
Planar reflective surfaces such as glass and mirrors are notoriously hard to reconstruct for most current 3D scanning techniques. When treated naïvely, they introduce duplicate scene structures, effectively destroying the reconstruction altogether. Our key insight is that an easy to identify structure attached to the scanner---in our case an AprilTag---can yield reliable information about the existence and the geometry of glass and mirror surfaces in a scene. We introduce a fully automatic pipeline that allows us to reconstruct the geometry and extent of planar glass and mirror surfaces while being able to distinguish between the two. Furthermore, our system can automatically segment observations of multiple reflective surfaces in a scene based on their estimated planes and locations. In the proposed setup, minimal additional hardware is needed to create high-quality results. We demonstrate this using reconstructions of several scenes with a variety of real mirrors and glass.
Thomas Whelan, Michael Goesele, Steven Lovegrove, Julian Straub, Simon Green, Richard Szeliski, Steven Butterfield, Shobhit Verma, Richard A. Newcombe
ACM Trans. Graph.6
2017 Video Segmentation with Background Motion Models
Scott Wehrwein, Richard Szeliski
BMVC2
2017 Casual 3D photography
abstract
We present an algorithm that enables casual 3D photography. Given a set of input photos captured with a hand-held cell phone or DSLR camera, our algorithm reconstructs a 3D photo , a central panoramic, textured, normal mapped, multi-layered geometric mesh representation. 3D photos can be stored compactly and are optimized for being rendered from viewpoints that are near the capture viewpoints. They can be rendered using a standard rasterization pipeline to produce perspective views with motion parallax. When viewed in VR, 3D photos provide geometrically consistent views for both eyes. Our geometric representation also allows interacting with the scene using 3D geometry-aware effects, such as adding new objects to the scene and artistic lighting effects. Our 3D photo reconstruction algorithm starts with a standard structure from motion and multi-view stereo reconstruction of the scene. The dense stereo reconstruction is made robust to the imperfect capture conditions using a novel near envelope cost volume prior that discards erroneous near depth hypotheses. We propose a novel parallax-tolerant stitching algorithm that warps the depth maps into the central panorama and stitches two color-and-depth panoramas for the front and back scene surfaces. The two panoramas are fused into a single non-redundant, well-connected geometric mesh. We provide videos demonstrating users interactively viewing and manipulating our 3D photos.
Peter Hedman, Suhib Alsisan, Richard Szeliski, Johannes Kopf 0001
ACM Trans. Graph.3
2017 Low-cost 360 stereo photography and video capture
abstract
A number of consumer-grade spherical cameras have recently appeared, enabling affordable monoscopic VR content creation in the form of full 360° X 180° spherical panoramic photos and videos. While monoscopic content is certainly engaging, it fails to leverage a main aspect of VR HMDs, namely stereoscopic display. Recent stereoscopic capture rigs involve placing many cameras in a ring and synthesizing an omni-directional stereo panorama enabling a user to look around to explore the scene in stereo. In this work, we describe a method that takes images from two 360° spherical cameras and synthesizes an omni-directional stereo panorama with stereo in all directions. Our proposed method has a lower equipment cost than camera-ring alternatives, can be assembled with currently available off-the-shelf equipment, and is relatively small and light-weight compared to the alternatives. We validate our method by generating both stills and videos. We have conducted a user study to better understand what kinds of geometric processing are necessary for a pleasant viewing experience. We also discuss several algorithmic variations, each with their own time and quality trade-offs.
Kevin Matzen, Michael F. Cohen, Bryce Evans, Johannes Kopf 0001, Richard Szeliski
ACM Trans. Graph.5
2016 Editorial
Jingyi Ju, Bastian Goldlücke, Richard Szeliski, Tomás Pajdla
Comput. Vis. Image Underst.3
2016 Guest Editorial: Special Section on CVPR 2013
abstract
This special section contains selected papers from the IEEE Computer Vision and Pattern Recognition (CVPR), June, 2013, jointly sponsored by the IEEE and the Computer Vision Foundation.
William T. Freeman, Richard Szeliski, Gregory D. Hager
IEEE Trans. Pattern Anal. Mach. Intell.2
2015 Light field layer matting
abstract
In this paper, we use matting to separate foreground layers from light fields captured with a plenoptic camera. We represent the input 4D light field as a 4D background light field, plus a 2D spatially varying foreground color layer with alpha. Our method can be used to both pull a foreground matte and estimate an occluded background light field. Our method assumes that the foreground layer is thin and fronto-parallel, and is composed of a limited set of colors that are distinct from the background layer colors. Our method works well for thin, translucent, and blurred foreground occluders. Our representation can be used to render the light field from novel views, handling disocclusions while avoiding common artifacts.
Juliet Fiss, Brian Curless, Richard Szeliski
CVPR3
2015 Model-Based Tracking at 300Hz Using Raw Time-of-Flight Observations
abstract
Consumer depth cameras have dramatically improved our ability to track rigid, articulated, and deformable 3D objects in real-time. However, depth cameras have a limited temporal resolution (frame-rate) that restricts the accuracy and robustness of tracking, especially for fast or unpredictable motion. In this paper, we show how to perform model-based object tracking which allows to reconstruct the object's depth at an order of magnitude higher frame-rate through simple modifications to an off-the-shelf depth camera. We focus on phase-based time-of-flight (ToF) sensing, which reconstructs each low frame-rate depth image from a set of short exposure 'raw' infrared captures. These raw captures are taken in quick succession near the beginning of each depth frame, and differ in the modulation of their active illumination. We make two contributions. First, we detail how to perform model-based tracking against these raw captures. Second, we show that by reprogramming the camera to space the raw captures uniformly in time, we obtain a 10x higher frame-rate, and thereby improve the ability to track fast-moving objects.
Jan Stühmer, Sebastian Nowozin, Andrew W. Fitzgibbon, Richard Szeliski, Travis Perry, Sunil Acharya, Daniel Cremers, Jamie Shotton
ICCV4
2014 Efficient High-Resolution Stereo Matching Using Local Plane Sweeps
abstract
We present a stereo algorithm designed for speed and efficiency that uses local slanted plane sweeps to propose disparity hypotheses for a semi-global matching algorithm. Our local plane hypotheses are derived from initial sparse feature correspondences followed by an iterative clustering step. Local plane sweeps are then performed around each slanted plane to produce out-of-plane parallax and matching-cost estimates. A final global optimization stage, implemented using semi-global matching, assigns each pixel to one of the local plane hypotheses. By only exploring a small fraction of the whole disparity space volume, our technique achieves significant speedups over previous algorithms and achieves state-of-the-art accuracy on high-resolution stereo pairs of up to 19 megapixels.
Sudipta N. Sinha, Daniel Scharstein, Richard Szeliski
CVPR3
2014 Refocusing plenoptic images using depth-adaptive splatting
abstract
In this paper, we propose a simple, novel plane sweep technique for refocusing plenoptic images. Rays are projected directly from the raw plenoptic image captured on the sensor into the output image plane, without computing intermediate representations such as subaperture views or epipolar images. Interpolation is performed in the output image plane using splatting. The splat kernel for each ray is adjusted adaptively, based on the refocus depth and an estimate of the depth at which that ray intersects the scene. This adaptive interpolation method antialiases out-of-focus regions, while keeping in-focus regions sharp. We test the proposed method on images from a Lytro camera and compare our results with those from the Lytro SDK. Additionally, we provide a thorough discussion of our calibration and preprocessing pipeline for this camera.
Juliet Fiss, Brian Curless, Richard Szeliski
ICCP3
2014 Car make and model recognition using 3D curve alignment
abstract
We present a new approach for recognizing the make and model of a car from a single image. While most previous methods are restricted to fixed or limited viewpoints, our system is able to verify a car's make and model from an arbitrary view. Our model consists of 3D space curves obtained by backprojecting image curves onto silhouette-based visual hulls and then refining them using three-view curve matching. We also build an appearance model of taillights which is used as an additional cue. Our approach is able to verify the exact make and model of a car over a wide range of viewpoints and background clutter.
Edward Hsiao, Sudipta N. Sinha, Krishnan Ramnath, Simon Baker, C. Lawrence Zitnick, Richard Szeliski
WACV6
2014 Car make and model recognition using 3D curve alignment
abstract
We present a new approach for recognizing the make and model of a car from a single image. While most previous methods are restricted to fixed or limited viewpoints, our system is able to verify a car's make and model from an arbitrary view. Our model consists of 3D space curves obtained by backprojecting image curves onto silhouette-based visual hulls and then refining them using three-view curve matching. These 3D curves are then matched to 2D image curves using a 3D view-based alignment technique. We present two different methods for estimating the pose of a car, which we then use to initialize the 3D curve matching. Our approach is able to verify the exact make and model of a car over a wide range of viewpoints in cluttered scenes.
Krishnan Ramnath, Sudipta N. Sinha, Richard Szeliski, Edward Hsiao
WACV3
2014 First-person hyper-lapse videos
abstract
We present a method for converting first-person videos, for example, captured with a helmet camera during activities such as rock climbing or bicycling, into hyper-lapse videos, i.e., time-lapse videos with a smoothly moving camera. At high speed-up rates, simple frame sub-sampling coupled with existing video stabilization methods does not work, because the erratic camera shake present in first-person videos is amplified by the speed-up. Our algorithm first reconstructs the 3D input camera path as well as dense, per-frame proxy geometries. We then optimize a novel camera path for the output video that passes near the input cameras while ensuring that the virtual camera looks in directions that can be rendered well from the input. Finally, we generate the novel smoothed, time-lapse video by rendering, stitching, and blending appropriately selected source frames for each output frame. We present a number of results for challenging videos that cannot be processed using traditional techniques.
Johannes Kopf 0001, Michael F. Cohen, Richard Szeliski
ACM Trans. Graph.3
2013 Image-based rendering in the gradient domain
abstract
We propose a novel image-based rendering algorithm for handling complex scenes that may include reflective surfaces. Our key contribution lies in treating the problem in the gradient domain. We use a standard technique to estimate scene depth, but assign depths to image gradients rather than pixels. A novel view is obtained by rendering the horizontal and vertical gradients, from which the final result is reconstructed through Poisson integration using an approximate solution as a data term. Our algorithm is able to handle general scenes including reflections and similar effects without explicitly separating the scene into reflective and transmissive parts, as required by previous work. Our prototype renderer is fully implemented on the GPU and runs in real time on commodity hardware.
Johannes Kopf 0001, Fabian Langguth, Daniel Scharstein, Richard Szeliski, Michael Goesele
ACM Trans. Graph.4
2013 Efficient preconditioning of laplacian matrices for computer graphics
abstract
We present a new multi-level preconditioning scheme for discrete Poisson equations that arise in various computer graphics applications such as colorization, edge-preserving decomposition for two-dimensional images, and geodesic distances and diffusion on three-dimensional meshes. Our approach interleaves the selection of fine-and coarse-level variables with the removal of weak connections between potential fine-level variables ( sparsification ) and the compensation for these changes by strengthening nearby connections. By applying these operations before each elimination step and repeating the procedure recursively on the resulting smaller systems, we obtain a highly efficient multi-level preconditioning scheme with linear time and memory requirements. Our experiments demonstrate that our new scheme outperforms or is comparable with other state-of-the-art methods, both in terms of operation count and wall-clock time. This speedup is achieved by the new method's ability to reduce the condition number of irregular Laplacian matrices as well as homogeneous systems. It can therefore be used for a wide variety of computational photography problems, as well as several 3D mesh processing tasks, without the need to carefully match the algorithm to the problem characteristics.
Dilip Krishnan, Raanan Fattal, Richard Szeliski
ACM Trans. Graph.3
2013 Navigating the worldwide community of photos
abstract
The last decade has seen an explosion in the number of photographs available on the Internet. The sheer volume of interesting photos makes it a challenge to explore this space. Various Web and social media sites, along with search and indexing techniques, have been developed in response. One natural way to navigate these images in a 3D geo-located context. In this article, we reflect on our work in this area, with a focus on techniques that build partial 3D scene models to help find and navigate interesting photographs in an interactive, immersive 3D setting. We also discuss how finding such relationships among photographs opens up exciting new possibilities for multimedia authoring, visualization, and editing.
Richard Szeliski, Noah Snavely, Steven M. Seitz
ACM Trans. Multim. Comput. Commun. Appl.1
2012 Multiple View Object Cosegmentation Using Appearance and Stereo Cues
Adarsh Kowdle, Sudipta N. Sinha, Richard Szeliski
ECCV (5)3
2012 Detecting and Reconstructing 3D Mirror Symmetric Objects
Sudipta N. Sinha, Krishnan Ramnath, Richard Szeliski
ECCV (2)3
2012 Image Restoration by Matching Gradient Distributions
abstract
The restoration of a blurry or noisy image is commonly performed with a MAP estimator, which maximizes a posterior probability to reconstruct a clean image from a degraded image. A MAP estimator, when used with a sparse gradient image prior, reconstructs piecewise smooth images and typically removes textures that are important for visual realism. We present an alternative deconvolution method called iterative distribution reweighting (IDR) which imposes a global constraint on gradients so that a reconstructed image should have a gradient distribution similar to a reference distribution. In natural images, a reference distribution not only varies from one image to another, but also within an image depending on texture. We estimate a reference distribution directly from an input image for each texture segment. Our algorithm is able to restore rich mid-frequency textures. A large-scale user study supports the conclusion that our algorithm improves the visual realism of reconstructed images compared to those of MAP estimators.
Taeg Sang Cho, C. Lawrence Zitnick, Neel Joshi, Sing Bing Kang, Richard Szeliski, William T. Freeman
IEEE Trans. Pattern Anal. Mach. Intell.5
2012 Pushing the Envelope of Modern Methods for Bundle Adjustment
abstract
In this paper, we present results and experiments with several methods for bundle adjustment, producing the fastest bundle adjuster ever published in terms of computation and convergence. From a computational perspective, the fastest methods naturally handle the block-sparse pattern that arises in a reduced camera system. Adapting to the naturally arising block-sparsity allows the use of BLAS3, efficient memory handling, fast variable ordering, and customized sparse solving, all simultaneously. We present two methods; one uses exact minimum degree ordering and block-based LDL solving and the other uses block-based preconditioned conjugate gradients. Both methods are performed on the reduced camera system. We show experimentally that the adaptation to the natural block sparsity allows both of these methods to perform better than previous methods. Further improvements in convergence speed are achieved by the novel use of embedded point iterations. Embedded point iterations take place inside each camera update step, yielding a greater cost decrease from each camera update step and, consequently, a lower minimum. This is especially true for points projecting far out on the flatter region of the robustifier. Intensive analyses from various angles demonstrate the improved performance of the presented bundler.
Yekeun Jeong, David Nistér, Drew Steedly, Richard Szeliski, In-So Kweon
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Image-based rendering for scenes with reflections
abstract
We present a system for image-based modeling and rendering of real-world scenes containing reflective and glossy surfaces. Previous approaches to image-based rendering assume that the scene can be approximated by 3D proxies that enable view interpolation using traditional back-to-front or z-buffer compositing. In this work, we show how these can be generalized to multiple layers that are combined in an additive fashion to model the reflection and transmission of light that occurs at specular surfaces such as glass and glossy materials. To simplify the analysis and rendering stages, we model the world using piecewise-planar layers combined using both additive and opaque mixing of light. We also introduce novel techniques for estimating multiple depths in the scene and separating the reflection and transmission components into different layers. We then use our system to model and render a variety of real-world scenes with reflections.
Sudipta N. Sinha, Johannes Kopf 0001, Michael Goesele, Daniel Scharstein, Richard Szeliski
ACM Trans. Graph.5
2011 Structure from motion for scenes with large duplicate structures
abstract
Most existing structure from motion (SFM) approaches for unordered images cannot handle multiple instances of the same structure in the scene. When image pairs containing different instances are matched based on visual similarity, the pairwise geometric relations as well as the correspondences inferred from such pairs are erroneous, which can lead to catastrophic failures in the reconstruction. In this paper, we investigate the geometric ambiguities caused by the presence of repeated or duplicate structures and show that to disambiguate between multiple hypotheses requires more than pure geometric reasoning. We couple an expectation maximization (EM)-based algorithm that estimates camera poses and identifies the false match-pairs with an efficient sampling method to discover plausible data association hypotheses. The sampling method is informed by geometric and image-based cues. Our algorithm usually recovers the correct data association, even in the presence of large numbers of false pairwise matches.
Richard Roberts 0001, Sudipta N. Sinha, Richard Szeliski, Drew Steedly
CVPR3
2011 Fast Poisson blending using multi-splines
abstract
We present a technique for fast Poisson blending and gradient domain compositing. Instead of using a single piecewise-smooth offset map to perform the blending, we associate a separate map with each input source image. Each individual offset map is itself smoothly varying and can therefore be represented using a low-dimensional spline. The resulting linear system is much smaller than either the original Poisson system or the quadtree spline approximation of a single (unified) offset map. We demonstrate the speed and memory improvements available with our system and apply it to large panoramas. We also show how robustly modeling the multiplicative gain rather than the offset between overlapping images leads to improved results, and how adding a small amount of Laplacian pyramid blending improves the results in areas of inconsistent texture.
Richard Szeliski, Matthew Uyttendaele, Drew Steedly
ICCP1
2011 A Database and Evaluation Methodology for Optical Flow
abstract
The quantitative evaluation of optical flow algorithms by Barron et al. (1994) led to significant advances in performance. The challenges for optical flow algorithms today go beyond the datasets and evaluation methods proposed in that paper. Instead, they center on problems associated with complex natural scenes, including nonrigid motion, real sensor noise, and motion discontinuities. We propose a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: (1) sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture, (2) realistic synthetic sequences, (3) high frame-rate video used to study interpolation error, and (4) modified stereo sequences of static scenes. In addition to the average angular error used by Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and results at motion discontinuities and in textureless regions. In October 2007, we published the performance of several well-known methods on a preliminary version of our data to establish the current state of the art. We also made the data freely available on the web at http://vision.middlebury.edu/flow/ . Subsequently a number of researchers have uploaded their results to our website and published papers using the data. A significant improvement in performance has already been achieved. In this paper we analyze the results obtained to date and draw a large number of conclusions from them.
Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski
Int. J. Comput. Vis.6
2011 Multigrid and multilevel preconditioners for computational photography
abstract
This paper unifies multigrid and multilevel (hierarchical) preconditioners, two widely-used approaches for solving computational photography and other computer graphics simulation problems. It provides detailed experimental comparisons of these techniques and their variants, including an analysis of relative computational costs and how these impact practical algorithm performance. We derive both theoretical convergence rates based on the condition numbers of the systems and their preconditioners, and empirical convergence rates drawn from real-world problems. We also develop new techniques for sparsifying higher connectivity problems, and compare our techniques to existing and newly developed variants such as algebraic and combinatorial multigrid. Our experimental results demonstrate that, except for highly irregular problems, adaptive hierarchical basis function preconditioners generally outperform alternative multigrid techniques, especially when computational complexity is taken into account.
Dilip Krishnan, Richard Szeliski
ACM Trans. Graph.2
2010 Removing rolling shutter wobble
abstract
We present an algorithm to remove wobble artifacts from a video captured with a rolling shutter camera undergoing large accelerations or jitter. We show how estimating the rapid motion of the camera can be posed as a temporal super-resolution problem. The low-frequency measurements are the motions of pixels from one frame to the next. These measurements are modeled as temporal integrals of the underlying high-frequency jitter of the camera. The estimated high-frequency motion of the camera is then used to re-render the sequence as though all the pixels in each frame were imaged at the same time. We also present an auto-calibration algorithm that can estimate the time between the capture of subsequent rows in the camera.
Simon Baker, Eric P. Bennett, Sing Bing Kang, Richard Szeliski
CVPR4
2010 A content-aware image prior
abstract
In image restoration tasks, a heavy-tailed gradient distribution of natural images has been extensively exploited as an image prior. Most image restoration algorithms impose a sparse gradient prior on the whole image, reconstructing an image with piecewise smooth characteristics. While the sparse gradient prior removes ringing and noise artifacts, it also tends to remove mid-frequency textures, degrading the visual quality. We can attribute such degradations to imposing an incorrect image prior. The gradient profile in fractal-like textures, such as trees, is close to a Gaussian distribution, and small gradients from such regions are severely penalized by the sparse gradient prior. To address this issue, we introduce an image restoration algorithm that adapts the image prior to the underlying texture. We adapt the prior to both low-level local structures as well as mid-level textural characteristics. Improvements in visual quality is demonstrated on deconvolution and denoising tasks.
Taeg Sang Cho, Neel Joshi, C. Lawrence Zitnick, Sing Bing Kang, Richard Szeliski, William T. Freeman
CVPR5
2010 Towards Internet-scale multi-view stereo
abstract
This paper introduces an approach for enabling existing multi-view stereo methods to operate on extremely large unstructured photo collections. The main idea is to decompose the collection into a set of overlapping sets of photos that can be processed in parallel, and to merge the resulting reconstructions. This overlapping clustering problem is formulated as a constrained optimization and solved iteratively. The merging algorithm, designed to be parallel and out-of-core, incorporates robust filtering steps to eliminate low-quality reconstructions and enforce global visibility constraints. The approach has been tested on several large datasets downloaded from Flickr.com, including one with over ten thousand images, yielding a 3D reconstruction with nearly thirty million points.
Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski
CVPR4
2010 Pushing the envelope of modern methods for bundle adjustment
abstract
In this paper, we present results and experiments with several methods for bundle adjustment, producing the fastest bundle adjuster ever published. The fastest methods work with the well known reduced camera system and handle the block-sparse pattern arising in the reduced camera system in a natural way. Adapting to the naturally arising block-sparsity allows the use of BLAS3, efficient memory handling, fast variable ordering, and customized sparse solving all at the same time. We present two methods, one using exact minimum degree ordering and block-based LDL solving, and one using block-based preconditioned conjugate gradient, both on the reduced camera system. We show experimentally that the adaptation to the natural block sparsity allows both these methods to perform better than previous ones. Further speed improvements are achieved by the novel use of embedded point iterations. The embedded point iterations take place inside each camera update step, yielding a higher cost decrease from each camera update step. This is especially true for points projecting far out on the flatter region of the robustifier.
Yekeun Jeong, David Nistér, Drew Steedly, Richard Szeliski, In-So Kweon
CVPR4
2010 Bundle Adjustment in the Large
Sameer Agarwal 0001, Noah Snavely, Steven M. Seitz, Richard Szeliski
ECCV (2)4
2010 Scene Reconstruction and Visualization From Community Photo Collections
abstract
There are billions of photographs on the Internet, representing an extremely large, rich, and nearly comprehensive visual record of virtually every famous place on Earth. Unfortunately, these massive community photo collections are almost completely unstructured, making it very difficult to use them for applications such as the virtual exploration of our world. Over the past several years, advances in computer vision have made it possible to automatically reconstruct 3-D geometry - including camera positions and scene models - from these large, diverse photo collections. Once the geometry is known, we can recover higher level information from the spatial distribution of photos, such as the most common viewpoints and paths through the scene. This paper reviews recent progress on these challenging computer vision problems, and describes how we can use the recovered structure to turn community photo collections into immersive, interactive 3-D experiences.
Noah Snavely, Ian Simon, Michael Goesele, Richard Szeliski, Steven M. Seitz
Proc. IEEE4
2010 Ambient point clouds for view interpolation
abstract
View interpolation and image-based rendering algorithms often produce visual artifacts in regions where the 3D scene geometry is erroneous, uncertain, or incomplete. We introduce ambient point clouds constructed from colored pixels with uncertain depth, which help reduce these artifacts while providing non-photorealistic background coloring and emphasizing reconstructed 3D geometry. Ambient point clouds are created by randomly sampling colored points along the viewing rays associated with uncertain pixels. Our real-time rendering system combines these with more traditional rigid 3D point clouds and colored surface meshes obtained using multiview stereo. Our resulting system can handle larger-range view transitions with fewer visible artifacts than previous approaches.
Michael Goesele, Jens Ackermann, Simon Fuhrmann, Carsten Haubold, Ronny Klowsky, Drew Steedly, Richard Szeliski
ACM Trans. Graph.7
2010 Image deblurring using inertial measurement sensors
abstract
We present a deblurring algorithm that uses a hardware attachment coupled with a natural image prior to deblur images from consumer cameras. Our approach uses a combination of inexpensive gyroscopes and accelerometers in an energy optimization framework to estimate a blur function from the camera's acceleration and angular velocity during an exposure. We solve for the camera motion at a high sampling rate during an exposure and infer the latent image using a joint optimization. Our method is completely automatic, handles per-pixel, spatially-varying blur, and out-performs the current leading image-based methods. Our experiments show that it handles large kernels -- up to at least 100 pixels, with a typical size of 30 pixels. We also present a method to perform "ground-truth" measurements of camera motion blur. We use this method to validate our hardware and deconvolution approach. To the best of our knowledge, this is the first work that uses 6 DOF inertial sensors for dense, per-pixel spatially-varying image deblurring and the first work to gather dense ground-truth measurements for camera-shake blur.
Neel Joshi, Sing Bing Kang, C. Lawrence Zitnick, Richard Szeliski
ACM Trans. Graph.4
2010 Street slide: browsing street level imagery
abstract
Systems such as Google Street View and Bing Maps Streetside enable users to virtually visit cities by navigating between immersive 360° panoramas, or bubbles. The discrete moves from bubble to bubble enabled in these systems do not provide a good visual sense of a larger aggregate such as a whole city block. Multi-perspective "strip" panoramas can provide a visual summary of a city street but lack the full realism of immersive panoramas. We present Street Slide, which combines the best aspects of the immersive nature of bubbles with the overview provided by multi-perspective strip panoramas. We demonstrate a seamless transition between bubbles and multi-perspective panoramas. We also present a dynamic construction of the panoramas which overcomes many of the limitations of previous systems. As the user slides sideways, the multi-perspective panorama is constructed and rendered dynamically to simulate either a perspective or hyper-perspective view. This provides a strong sense of parallax, which adds to the immersion. We call this form of sliding sideways while looking at a street façade a street slide. Finally we integrate annotations and a mini-map within the user interface to provide geographic information as well additional affordances for navigation. We demonstrate our Street Slide system on a series of intersecting streets in an urban setting. We report the results of a user study, which shows that visual searching is greatly enhanced with the Street Slide interface over existing systems from Google and Bing.
Johannes Kopf 0001, Billy Chen, Richard Szeliski, Michael F. Cohen
ACM Trans. Graph.3
2009 Manhattan-world stereo
abstract
Multi-view stereo (MVS) algorithms now produce reconstructions that rival laser range scanner accuracy. However, stereo algorithms require textured surfaces, and therefore work poorly for many architectural scenes (e.g., building interiors with textureless, painted walls). This paper presents a novel MVS approach to overcome these limitations for Manhattan World scenes, i.e., scenes that consists of piece-wise planar surfaces with dominant directions. Given a set of calibrated photographs, we first reconstruct textured regions using an existing MVS algorithm, then extract dominant plane directions, generate plane hypotheses, and recover per-view depth maps using Markov random fields. We have tested our algorithm on several datasets ranging from office interiors to outdoor buildings, and demonstrate results that outperform the current state of the art for such texture-poor scenes.
Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski
CVPR4
2009 Image deblurring and denoising using color priors
abstract
Image blur and noise are difficult to avoid in many situations and can often ruin a photograph. We present a novel image deconvolution algorithm that deblurs and denoises an image given a known shift-invariant blur kernel. Our algorithm uses local color statistics derived from the image as a constraint in a unified framework that can be used for deblurring, denoising, and upsampling. A pixel's color is required to be a linear combination of the two most prevalent colors within a neighborhood of the pixel. This two-color prior has two major benefits: it is tuned to the content of the particular image and it serves to decouple edge sharpness from edge strength. Our unified algorithm for deblurring and denoising out-performs previous methods that are specialized for these individual applications. We demonstrate this with both qualitative results and extensive quantitative comparisons that show that we can out-perform previous methods by approximately 1 to 3 DB.
Neel Joshi, C. Lawrence Zitnick, Richard Szeliski, David J. Kriegman
CVPR3
2009 Building Rome in a day
abstract
We present a system that can match and reconstruct 3D scenes from extremely large collections of photographs such as those found by searching for a given city (e.g., Rome) on Internet photo sharing sites. Our system uses a collection of novel parallel distributed matching and reconstruction algorithms, designed to maximize parallelism at each stage in the pipeline and minimize serialization bottlenecks. It is designed to scale gracefully with both the size of the problem and the amount of available computation. We have experimented with a variety of alternative algorithms at each stage of the pipeline and report on which ones work best in a parallel computing environment. Our experimental results demonstrate that it is now possible to reconstruct cities consisting of 150 K images in less than a day on a cluster with 500 compute cores.
Sameer Agarwal 0001, Noah Snavely, Ian Simon, Steven M. Seitz, Richard Szeliski
ICCV5
2009 Reconstructing building interiors from images
abstract
This paper proposes a fully automated 3D reconstruction and visualization system for architectural scenes (interiors and exteriors). The reconstruction of indoor environments from photographs is particularly challenging due to texture-poor planar surfaces such as uniformly-painted walls. Our system first uses structure-from-motion, multi-view stereo, and a stereo algorithm specifically designed for Manhattan-world scenes (scenes consisting predominantly of piece-wise planar surfaces with dominant directions) to calibrate the cameras and to recover initial 3D geometry in the form of oriented points and depth maps. Next, the initial geometry is fused into a 3D model with a novel depth-map integration algorithm that, again, makes use of Manhattan-world assumptions and produces simplified 3D models. Finally, the system enables the exploration of reconstructed environments with an interactive, image-based 3D viewer. We demonstrate results on several challenging datasets, including a 3D reconstruction and image-based walk-through of an entire floor of a house, the first result of this kind from an automated computer vision system.
Yasutaka Furukawa, Brian Curless, Steven M. Seitz, Richard Szeliski
ICCV4
2009 Piecewise planar stereo for image-based rendering
abstract
We present a novel multi-view stereo method designed for image-based rendering that generates piecewise planar depth maps from an unordered collection of photographs.
Sudipta N. Sinha, Drew Steedly, Richard Szeliski
ICCV3
2008 PSF estimation using sharp edge prediction
abstract
Image blur is caused by a number of factors such as motion, defocus, capturing light over the non-zero area of the aperture and pixel, the presence of anti-aliasing filters on a camera sensor, and limited sensor resolution. We present an algorithm that estimates non-parametric, spatially-varying blur functions (i.e., point-spread functions or PSFs) at subpixel resolution from a single image. Our method handles blur due to defocus, slight camera motion, and inherent aspects of the imaging system. Our algorithm can be used to measure blur due to limited sensor resolution by estimating a sub-pixel, super-resolved PSF even for in-focus images. It operates by predicting a ldquosharprdquo version of a blurry input image and uses the two images to solve for a PSF. We handle the cases where the scene content is unknown and also where a known printed calibration target is placed in the scene. Our method is completely automatic, fast, and produces accurate results.
Neel Joshi, Richard Szeliski, David J. Kriegman
CVPR2
2008 Skeletal graphs for efficient structure from motion
abstract
We address the problem of efficient structure from motion for large, unordered, highly redundant, and irregularly sampled photo collections, such as those found on Internet photo-sharing sites. Our approach computes a small skeletal subset of images, reconstructs the skeletal set, and adds the remaining images using pose estimation. Our technique drastically reduces the number of parameters that are considered, resulting in dramatic speedups, while provably approximating the covariance of the full set of parameters. To compute a skeletal image set, we first estimate the accuracy of two-frame reconstructions between pairs of overlapping images, then use a graph algorithm to select a subset of images that, when reconstructed, approximates the accuracy of the full set. A final bundle adjustment can then optionally be used to restore any loss of accuracy.
Noah Snavely, Steven M. Seitz, Richard Szeliski
CVPR3
2008 Modeling the World from Internet Photo Collections
Noah Snavely, Steven M. Seitz, Richard Szeliski
Int. J. Comput. Vis.3
2008 Automatic Estimation and Removal of Noise from a Single Image
abstract
Image denoising algorithms often assume an additive white Gaussian noise (AWGN) process that is independent of the actual RGB values. Such approaches are not fully automatic and cannot effectively remove color noise produced by todays CCD digital camera. In this paper, we propose a unified framework for two tasks: automatic estimation and removal of color noise from a single image using piecewise smooth image models. We introduce the noise level function (NLF), which is a continuous function describing the noise level as a function of image brightness. We then estimate an upper bound of the real noise level function by fitting a lower envelope to the standard deviations of per-segment image variances. For denoising, the chrominance of color noise is significantly removed by projecting pixel values onto a line fit to the RGB values in each segment. Then, a Gaussian conditional random field (GCRF) is constructed to obtain the underlying clean image from the noisy input. Extensive experiments are conducted to test the proposed algorithm, which is shown to outperform state-of-the-art denoising algorithms.
Ce Liu 0001, Richard Szeliski, Sing Bing Kang, C. Lawrence Zitnick, William T. Freeman
IEEE Trans. Pattern Anal. Mach. Intell.2
2008 A Comparative Study of Energy Minimization Methods for Markov Random Fields with Smoothness-Based Priors
abstract
Among the most exciting advances in early vision has been the development of efficient energy minimization algorithms for pixel-labeling tasks such as depth or texture computation. It has been known for decades that such problems can be elegantly expressed as Markov random fields, yet the resulting energy minimization problems have been widely viewed as intractable. Recently, algorithms such as graph cuts and loopy belief propagation (LBP) have proven to be very powerful: for example, such methods form the basis for almost all the top-performing stereo methods. However, the tradeoffs among different energy minimization algorithms are still not well understood. In this paper we describe a set of energy minimization benchmarks and use them to compare the solution quality and running time of several common energy minimization algorithms. We investigate three promising recent methods graph cuts, LBP, and tree-reweighted message passing in addition to the well-known older iterated conditional modes (ICM) algorithm. Our benchmark problems are drawn from published energy functions used for stereo, image stitching, interactive segmentation, and denoising. We also provide a general-purpose software interface that allows vision researchers to easily switch between optimization methods. Benchmarks, code, images, and results are available at http://vision.middlebury.edu/MRF/.
Richard Szeliski, Ramin Zabih, Daniel Scharstein, Olga Veksler, Vladimir Kolmogorov, Aseem Agarwala, Marshall F. Tappen, Carsten Rother
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Edge-preserving decompositions for multi-scale tone and detail manipulation
abstract
Many recent computational photography techniques decompose an image into a piecewise smooth base layer, containing large scale variations in intensity, and a residual detail layer capturing the smaller scale details in the image. In many of these applications, it is important to control the spatial scale of the extracted details, and it is often desirable to manipulate details at multiple scales, while avoiding visual artifacts. In this paper we introduce a new way to construct edge-preserving multi-scale image decompositions. We show that current basedetail decomposition techniques, based on the bilateral filter, are limited in their ability to extract detail at arbitrary scales. Instead, we advocate the use of an alternative edge-preserving smoothing operator, based on the weighted least squares optimization framework, which is particularly well suited for progressive coarsening of images and for multi-scale detail extraction. After describing this operator, we show how to use it to construct edge-preserving multi-scale decompositions, and compare it to the bilateral filter, as well as to other schemes. Finally, we demonstrate the effectiveness of our edge-preserving decompositions in the context of LDR and HDR tone mapping, detail enhancement, and other applications.
Zeev Farbman, Raanan Fattal, Dani Lischinski, Richard Szeliski
ACM Trans. Graph.4
2008 Interactive 3D architectural modeling from unordered photo collections
abstract
We present an interactive system for generating photorealistic, textured, piecewise-planar 3D models of architectural structures and urban scenes from unordered sets of photographs. To reconstruct 3D geometry in our system, the user draws outlines overlaid on 2D photographs. The 3D structure is then automatically computed by combining the 2D interaction with the multi-view geometric information recovered by performing structure from motion analysis on the input photographs. We utilize vanishing point constraints at multiple stages during the reconstruction, which is particularly useful for architectural scenes where parallel lines are abundant. Our approach enables us to accurately model polygonal faces from 2D interactions in a single image. Our system also supports useful operations such as edge snapping and extrusions. Seamless texture maps are automatically generated by combining multiple input photographs using graph cut optimization and Poisson blending. The user can add brush strokes as hints during the texture generation stage to remove artifacts caused by unmodeled geometric structures. We build models for a variety of architectural scenes from collections of up to about a hundred photographs.
Sudipta N. Sinha, Drew Steedly, Richard Szeliski, Maneesh Agrawala, Marc Pollefeys
ACM Trans. Graph.3
2008 Finding paths through the world's photos
abstract
When a scene is photographed many times by different people, the viewpoints often cluster along certain paths. These paths are largely specific to the scene being photographed, and follow interesting regions and viewpoints. We seek to discover a range of such paths and turn them into controls for image-based rendering. Our approach takes as input a large set of community or personal photos, reconstructs camera viewpoints, and automatically computes orbits, panoramas, canonical views, and optimal paths between views. The scene can then be interactively browsed in 3D using these controls or with six degree-of-freedom free-viewpoint control. As the user browses the scene, nearby views are continuously selected and transformed, using control-adaptive reprojection techniques.
Noah Snavely, Rahul Garg 0002, Steven M. Seitz, Richard Szeliski
ACM Trans. Graph.4
2007 City-Scale Location Recognition
abstract
We look at the problem of location recognition in a large image dataset using a vocabulary tree. This entails finding the location of a query image in a large dataset containing 3times104streetside images of a city. We investigate how the traditional invariant feature matching approach falls down as the size of the database grows. In particular we show that by carefully selecting the vocabulary using the most informative features, retrieval performance is significantly improved, allowing us to increase the number of database images by a factor of 10. We also introduce a generalization of the traditional vocabulary tree search algorithm which improves performance by effectively increasing the branching factor of a fixed vocabulary tree.
Grant Schindler, Matthew A. Brown, Richard Szeliski
CVPR3
2007 Layered Depth Panoramas
abstract
Representations for interactive photorealistic visualization of scenes range from compact 2D panoramas to data-intensive 4D light fields. In this paper, we propose a technique for creating a layered representation from a sparse set of images taken with a hand-held camera. This representation, which we call a layered depth panorama (LDP), allows the user to experience 3D by off-axis panning. It combines the compelling experience of panoramas with limited 3D navigation. Our choice of representation is motivated by ease of capture and compactness. We formulate the problem of constructing the LDP as the recovery of color and geometry in a multi-perspective cylindrical disparity space. We leverage a graph cut approach to sequentially determine the disparity and color of each layer using multi-view stereo. Geometry visible through the cracks at depth discontinuities in a frontmost layer is determined and assigned to layers behind the frontmost layer. All layers are then used to render novel panoramic views with parallax. We demonstrate our approach on a variety of complex outdoor and indoor scenes.
Ke Colin Zheng, Sing Bing Kang, Michael F. Cohen, Richard Szeliski
CVPR4
2007 A Database and Evaluation Methodology for Optical Flow
abstract
The quantitative evaluation of optical flow algorithms by Barron et al. led to significant advances in the performance of optical flow methods. The challenges for optical flow today go beyond the datasets and evaluation methods proposed in that paper and center on problems associated with nonrigid motion, real sensor noise, complex natural scenes, and motion discontinuities. Our goal is to establish a new set of benchmarks and evaluation methods for the next generation of optical flow algorithms. To that end, we contribute four types of data to test different aspects of optical flow algorithms: sequences with nonrigid motion where the ground-truth flow is determined by tracking hidden fluorescent texture; realistic synthetic sequences; high frame-rate video used to study interpolation error; and modified stereo sequences of static scenes. In addition to the average angular error used in Barron et al., we compute the absolute flow endpoint error, measures for frame interpolation error, improved statistics, and flow accuracy at motion boundaries and in textureless regions. We evaluate the performance of several well-known methods on this data to establish the current state of the art. Our database is freely available on the web together with scripts for scoring and publication of the results at http://vision.middlebury.edu/flow/.
Simon Baker, Daniel Scharstein, John P. Lewis, Stefan Roth 0001, Michael J. Black, Richard Szeliski
ICCV6
2007 Editorial
Katsushi Ikeuchi, Gudrun Klinker, Yuichi Ohta, Richard Szeliski
Int. J. Comput. Vis.4
2006 Finding People in Repeated Shots of the Same Scene
abstract
The goal of this work is to find all occurrences of a particular person in a sequence of photographs taken over a short period of time. For identification, we assume each individual’s hair and clothing stays the same throughout the sequence. Even with these assumptions, the task remains challenging as people can move around, change their pose and scale, and partially occlude each other. We propose a two stage method. First, individuals are identified by clustering frontal face detections using color clothing information. Second, a color based pictorial structure model is used to find occurrences of each person in images where their frontal face detection was missed. Two extensions improving the pictorial structure detections are also described. In the first extension, we obtain a better clothing segmentation to improve the accuracy of the clothing color model. In the second extension, we simultaneously consider multiple detection hypotheses of all people potentially present in the shot. Our results show that people can be re-detected in images where they do not face the camera. Results are presented on several sequences from a personal photo collection.
Josef Sivic, C. Lawrence Zitnick, Richard Szeliski
BMVC3
2006 Seamless Image Stitching of Scenes with Large Motions and Exposure Differences
abstract
This paper presents a technique to automatically stitch multiple images at varying orientations and exposures to create a composite panorama that preserves the angular extent and dynamic range of the inputs. The main contribution of our method is that it allows for large exposure differences, large scene motion or other misregistrations between frames and requires no extra camera hardware. To do this, we introduce a two-step graph cut approach. The purpose of the first step is to fix the positions of moving objects in the scene. In the second step, we fill in the entire available dynamic range. We introduce data costs that encourage consistency and higher signal-to-noise ratios, and seam costs that encourage smooth transitions. Our method is simple to implement and effective. We demonstrate the effectiveness of our approach on several input sets with varying exposures and camera orientations.
Ashley Eden, Matthew Uyttendaele, Richard Szeliski
CVPR (2)3
2006 Noise Estimation from a Single Image
abstract
In order to work well, many computer vision algorithms require that their parameters be adjusted according to the image noise level, making it an important quantity to estimate. We show how to estimate an upper bound on the noise level from a single image based on a piecewise smooth image prior model and measured CCD camera response functions. We also learn the space of noise level functions how noise level changes with respect to brightness and use Bayesian MAP inference to infer the noise level function from a single image. We illustrate the utility of this noise estimation for two algorithms: edge detection and featurepreserving smoothing through bilateral filtering. For a variety of different noise levels, we obtain good results for both these algorithms with no user-specified inputs.
Ce Liu 0001, William T. Freeman, Richard Szeliski, Sing Bing Kang
CVPR (1)3
2006 A Comparison and Evaluation of Multi-View Stereo Reconstruction Algorithms
abstract
This paper presents a quantitative comparison of several multi-view stereo reconstruction algorithms. Until now, the lack of suitable calibrated multi-view image datasets with known ground truth (3D shape models) has prevented such direct comparisons. In this paper, we first survey multi-view stereo algorithms and compare them qualitatively using a taxonomy that differentiates their key properties. We then describe our process for acquiring and calibrating multiview image datasets with high-accuracy ground truth and introduce our evaluation methodology. Finally, we present the results of our quantitative comparison of state-of-the-art multi-view stereo reconstruction algorithms on six benchmark datasets. The datasets, evaluation details, and instructions for submitting new models are available online at http://vision.middlebury.edu/mview.
Steven M. Seitz, Brian Curless, James Diebel, Daniel Scharstein, Richard Szeliski
CVPR (1)5
2006 Reconstructing Occluded Surfaces Using Synthetic Apertures: Stereo, Focus and Robust Measures
abstract
Most algorithms for 3D reconstruction from images use cost functions based on SSD, which assume that the surfaces being reconstructed are visible to all cameras. This makes it difficult to reconstruct objects which are partially occluded. Recently, researchers working with large camera arrays have shown it is possible to "see through" occlusions using a technique called synthetic aperture focusing. This suggests that we can design alternative cost functions that are robust to occlusions using synthetic apertures. Our paper explores this design space. We compare classical shape from stereo with shape from synthetic aperture focus. We also describe two variants of multi-view stereo based on color medians and entropy that increase robustness to occlusions. We present an experimental comparison of these cost functions on complex light fields, measuring their accuracy against the amount of occlusion.
Vaibhav Vaish, Marc Levoy, Richard Szeliski, C. Lawrence Zitnick, Sing Bing Kang
CVPR (2)3
2006 Video and Image Bayesian Demosaicing with a Two Color Image Prior
Eric P. Bennett, Matthew Uyttendaele, C. Lawrence Zitnick, Richard Szeliski, Sing Bing Kang
ECCV (1)4
2006 A Comparative Study of Energy Minimization Methods for Markov Random Fields
Richard Szeliski, Ramin Zabih, Daniel Scharstein, Olga Veksler, Vladimir Kolmogorov, Aseem Agarwala, Marshall F. Tappen, Carsten Rother
ECCV (2)1
2006 Boundary matting for view synthesis
Samuel W. Hasinoff, Sing Bing Kang, Richard Szeliski
Comput. Vis. Image Underst.3
2006 Stereo Matching with Linear Superposition of Layers
abstract
In this paper, we address stereo matching in the presence of a class of non-Lambertian effects, where image formation can be modeled as the additive superposition of layers at different depths. The presence of such effects makes it impossible for traditional stereo vision algorithms to recover depths using direct color matching-based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, matching is accomplished using spatial-temporal differencing. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm which we proved to be convergent. Our algorithm recovers depth and color estimates for both synthetic and real image sequences.
Yanghai Tsin, Sing Bing Kang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.3
2006 Photographing long scenes with multi-viewpoint panoramas
abstract
We present a system for producing multi-viewpoint panoramas of long, roughly planar scenes, such as the facades of buildings along a city street, from a relatively sparse set of photographs captured with a handheld still camera that is moved along the scene. Our work is a significant departure from previous methods for creating multi-viewpoint panoramas, which composite thin vertical strips from a video sequence captured by a translating video camera, in that the resulting panoramas are composed of relatively large regions of ordinary perspective. In our system, the only user input required beyond capturing the photographs themselves is to identify the dominant plane of the photographed scene; our system then computes a panorama automatically using Markov Random Field optimization. Users may exert additional control over the appearance of the result by drawing rough strokes that indicate various high-level goals. We demonstrate the results of our system on several scenes, including urban streets, a river bank, and a grocery store aisle.
Aseem Agarwala, Maneesh Agrawala, Michael F. Cohen, David Salesin, Richard Szeliski
ACM Trans. Graph.5
2006 Interactive local adjustment of tonal values
abstract
This paper presents a new interactive tool for making local adjustments of tonal values and other visual parameters in an image. Rather than carefully selecting regions or hand-painting layer masks, the user quickly indicates regions of interest by drawing a few simple brush strokes and then uses sliders to adjust the brightness, contrast, and other parameters in these regions. The effects of the user's sparse set of constraints are interpolated to the entire image using an edge-preserving energy minimization method designed to prevent the propagation of tonal adjustments to regions of significantly different luminance. The resulting system is suitable for adjusting ordinary and high dynamic range images, and provides the user with much more creative control than existing tone mapping algorithms. Our tool is also able to produce a tone mapping automatically, which may serve as a basis for further local adjustments, if so desired. The constraint propagation approach developed in this paper is a general one, and may also be used to interactively control a variety of other adjustments commonly performed in the digital darkroom.
Dani Lischinski, Zeev Farbman, Matthew Uyttendaele, Richard Szeliski
ACM Trans. Graph.4
2006 Photo tourism: exploring photo collections in 3D
abstract
We present a system for interactively browsing and exploring large unstructured collections of photographs of a scene using a novel 3D interface. Our system consists of an image-based modeling front end that automatically computes the viewpoint of each photograph as well as a sparse 3D model of the scene and image to model correspondences. Our photo explorer uses image-based rendering techniques to smoothly transition between photographs, while also enabling full 3D navigation and exploration of the set of images and world geometry, along with auxiliary information such as overhead maps. Our system also makes it easy to construct photo tours of scenic or historic locations, and to annotate image details, which are automatically transferred to other relevant images. We demonstrate our system on several large personal photo collections as well as images gathered from Internet photo sharing sites.
Noah Snavely, Steven M. Seitz, Richard Szeliski
ACM Trans. Graph.3
2006 Locally adapted hierarchical basis preconditioning
abstract
This paper develops locally adapted hierarchical basis functions for effectively preconditioning large optimization problems that arise in computer graphics applications such as tone mapping, gradient-domain blending, colorization, and scattered data interpolation. By looking at the local structure of the coefficient matrix and performing a recursive set of variable eliminations, combined with a simplification of the resulting coarse level problems, we obtain bases better suited for problems with inhomogeneous (spatially varying) data, smoothness, and boundary constraints. Our approach removes the need to heuristically adjust the optimal number of preconditioning levels, significantly outperforms previously proposed approaches, and also maps cleanly onto data-parallel architectures such as modern GPUs.
Richard Szeliski
ACM Trans. Graph.1
2005 Multi-Image Matching Using Multi-Scale Oriented Patches
abstract
This paper describes a novel multi-view matching framework based on a new type of invariant feature. Our features are located at Harris corners in discrete scale-space and oriented using a blurred local gradient. This defines a rotationally invariant frame in which we sample a feature descriptor, which consists of an 8 /spl times/ 8 patch of bias/gain normalised intensity values. The density of features in the image is controlled using a novel adaptive non-maximal suppression algorithm, which gives a better spatial distribution of features than previous approaches. Matching is achieved using a fast nearest neighbour algorithm that indexes features based on their low frequency Haar wavelet coefficients. We also introduce a novel outlier rejection procedure that verifies a pairwise feature match based on a background distribution of incorrect feature matches. Feature matches are refined using RANSAC and used in an automatic 2D panorama stitcher that has been extensively tested on hundreds of sample inputs.
Matthew A. Brown, Richard Szeliski, Simon A. J. Winder
CVPR (1)2
2005 Efficiently Registering Video into Panoramic Mosaics
abstract
We present an automatic and efficient method to register and stitch thousands of video frames into a large panoramic mosaic. Our method preserves the robustness and accuracy of image stitchers that match all pairs of images while utilizing the ordering information provided by video. We reduce the cost of searching for matches between video frames by adaptively identifying key frames based on the amount of image-to-image overlap. Key frames are matched to all other key frames, but intermediate video frames are only matched to temporally neighboring key frames and intermediate frames. Image orientations can be estimated from this sparse set of matches in time quadratic to cubic in the number of key frames but only linear in the number of intermediate frames. Additionally, the matches between pairs of images are compressed by replacing measurements within small windows in the image with a single representative measurement. We show that this approach substantially reduces the time required to estimate the image orientations with minimal loss of accuracy. Finally, we demonstrate both the efficiency and quality of our results by registering several long video sequences
Drew Steedly, Christopher Joseph Pal, Richard Szeliski
ICCV3
2005 Extracting layers and analyzing their specular properties using epipolar-plane-image analysis
Antonio Criminisi, Sing Bing Kang, Rahul Swaminathan, Richard Szeliski, P. Anandan 0001
Comput. Vis. Image Underst.4
2005 Panoramic video textures
abstract
This paper describes a mostly automatic method for taking the output of a single panning video camera and creating a panoramic video texture (PVT): a video that has been stitched into a single, wide field of view and that appears to play continuously and indefinitely. The key problem in creating a PVT is that although only a portion of the scene has been imaged at any given time, the output must simultaneously portray motion throughout the scene. Like previous work in video textures, our method employs min-cut optimization to select fragments of video that can be stitched together both spatially and temporally. However, it differs from earlier work in that the optimization must take place over a much larger set of data. Thus, to create PVTs, we introduce a dynamic programming step, followed by a novel hierarchical min-cut optimization algorithm. We also use gradient-domain compositing to further smooth boundaries between video fragments. We demonstrate our results with an interactive viewer in which users can interactively pan and zoom on high-resolution PVTs.
Aseem Agarwala, Ke Colin Zheng, Christopher Joseph Pal, Maneesh Agrawala, Michael F. Cohen, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.8
2005 Animating pictures with stochastic motion textures
abstract
In this paper, we explore the problem of enhancing still pictures with subtly animated motions. We limit our domain to scenes containing passive elements that respond to natural forces in some fashion. We use a semi-automatic approach, in which a human user segments the scene into a series of layers to be individually animated. Then, a "stochastic motion texture" is automatically synthesized using a spectral method, i.e., the inverse Fourier transform of a filtered noise spectrum. The motion texture is a time-varying 2D displacement map, which is applied to each layer. The resulting warped layers are then recomposited to form the animated frames. The result is a looping video texture created from a single still image, which has the advantages of being more controllable and of generally higher image quality and resolution than a video texture created from a video source. We demonstrate the technique on a variety of photographs and paintings.
Yung-Yu Chuang, Dan B. Goldman, Ke Colin Zheng, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.6
2004 Visual Odometry and Map Correlation
Anat Levin, Richard Szeliski
CVPR (1)2
2004 Probability Models for High Dynamic Range Imaging
Christopher Joseph Pal, Richard Szeliski, Matthew Uyttendaele, Nebojsa Jojic
CVPR (2)2
2004 Extracting View-Dependent Depth Maps from a Collection of Images
Sing Bing Kang, Richard Szeliski
Int. J. Comput. Vis.2
2004 Stereo Reconstruction from Multiperspective Panoramas
Yin Li 0003, Harry Shum, Chi-Keung Tang, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.4
2004 Sampling the Disparity Space Image
abstract
A central issue in stereo algorithm design is the choice of matching cost. Many algorithms simply use squared or absolute intensity differences based on integer disparity steps. In this paper, we address potential problems with such approaches. We begin with a careful analysis of the properties of the continuous disparity space image (DSI) and propose several new matching cost variants based on symmetrically matching interpolated image signals. Using stereo images with ground truth, we empirically evaluate the performance of the different cost variants and show that proper sampling can yield improved matching performance.
Richard Szeliski, Daniel Scharstein
IEEE Trans. Pattern Anal. Mach. Intell.1
2004 Digital photography with flash and no-flash image pairs
abstract
Digital photography has made it possible to quickly and easily take a pair of images of low-light environments: one with flash to capture detail and one without flash to capture ambient illumination. We present a variety of applications that analyze and combine the strengths of such flash/no-flash image pairs. Our applications include denoising and detail transfer (to merge the ambient qualities of the no-flash image with the high-frequency flash detail), white-balancing (to change the color tone of the ambient image), continuous flash (to interactively adjust flash intensity), and red-eye removal (to repair artifacts in the flash image). We demonstrate how these applications can synthesize new images that are of higher quality than either of the originals.
Georg Petschnigg, Richard Szeliski, Maneesh Agrawala, Michael F. Cohen, Hugues Hoppe, Kentaro Toyama
ACM Trans. Graph.2
2004 High-quality video view interpolation using a layered representation
abstract
The ability to interactively control viewpoint while watching a video is an exciting application of image-based rendering. The goal of our work is to render dynamic scenes with interactive viewpoint control using a relatively small number of video cameras. In this paper, we show how high-quality video-based rendering of dynamic scenes can be accomplished using multiple synchronized video streams combined with novel image-based modeling and rendering algorithms. Once these video streams have been processed, we can synthesize any intermediate view between cameras at any time, with the potential for space-time manipulation.In our approach, we first use a novel color segmentation-based stereo algorithm to generate high-quality photoconsistent correspondences across all camera views. Mattes for areas near depth discontinuities are then automatically extracted to reduce artifacts during view synthesis. Finally, a novel temporal two-layer compressed representation that handles matting is developed for rendering at interactive rates.
C. Lawrence Zitnick, Sing Bing Kang, Matthew Uyttendaele, Simon A. J. Winder, Richard Szeliski
ACM Trans. Graph.5
2003 High-Accuracy Stereo Depth Maps Using Structured Light
abstract
Progress in stereo algorithm performance is quickly outpacing the ability of existing stereo data sets to discriminate among the best-performing algorithms, motivating the need for more challenging scenes with accurate ground truth information. This paper describes a method for acquiring high-complexity stereo image pairs with pixel-accurate correspondence information using structured light. Unlike traditional range-sensing approaches, our method does not require the calibration of the light sources and yields registered disparity maps between all pairs of cameras and illumination projectors. We present new stereo data sets acquired with our method and demonstrate their suitability for stereo algorithm evaluation. Our results are available at http://www.middlebury.edu/stereo/.
Daniel Scharstein, Richard Szeliski
CVPR (1)2
2003 Stereo Matching with Reflections and Translucency
abstract
In this paper, we address the stereo matching problem in the presence of reflections and translucency, where image formation can be modeled as the additive superposition of layers at different depth. The presence of such effects violates the Lambertian assumption underlying traditional stereo vision algorithms, making it impossible to recover component depths using direct color matching based methods. We develop several techniques to estimate both depths and colors of the component layers. Depth hypotheses are enumerated in pairs, one from each layer, in a nested plane sweep. For each pair of depth hypotheses, we compute a component-color-independent matching error per pixel, using a spatial-temporal differencing technique. We then use graph cut optimization to solve for the depths of both layers. This is followed by an iterative color update algorithm whose convergence is proven in our paper. We show convincing results of depth and color estimates for both synthetic and real image sequences.
Yanghai Tsin, Sing Bing Kang, Richard Szeliski
CVPR (1)3
2003 Using Character Recognition and Segmentation to Tell Computer from Humans
abstract
How do you tell a computer from a human? The situation arises often on the Internet, when online polls are conducted, accounts are requested, undesired email is received, and chat-rooms are spammed. The approach we use is to create a visual challenge that is easy for humans but difficult for a computer. More specifically, our challenge is to recognize a string of random distorted characters. To pass the challenge, the subject must type in the correct corresponding ASCII string. From an OCR point of view, this problem is interesting because our goal is to use the vast amount of accumulated knowledge to defeat the state of the art OCR algorithms. This is a role reversal from traditional OCR research. Unlike many other systems, our algorithm is based on the assumption that segmentation is much more difficult than recognition. Our image challenges present hard segmentation problems that humans are particularly apt at solving. The technology is currently being used in MSN's Hotmail registration system, where it has significantly reduced daily registration rate with minimal Consumer Support impact.
Patrice Y. Simard, Richard Szeliski, Josh Benaloh, Julien Couvreur, Iulian Calinov
ICDAR2
2003 Shadow matting and compositing
abstract
In this paper, we describe a method for extracting shadows from one natural scene and inserting them into another. We develop physically-based shadow matting and compositing equations and use these to pull a shadow matte from a source scene in which the shadow is cast onto an arbitrary planar background. We then acquire the photometric and geometric properties of the target scene by sweeping oriented linear shadows (cast by a straight object) across it. From these shadow scans, we can construct a shadow displacement map without requiring camera or light source calibration. This map can then be used to deform the original shadow matte. We demonstrate our approach for both indoor scenes with controlled lighting and for outdoor scenes using natural lighting.
Yung-Yu Chuang, Dan B. Goldman, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.5
2003 High dynamic range video
abstract
Typical video footage captured using an off-the-shelf camcorder suffers from limited dynamic range. This paper describes our approach to generate high dynamic range (HDR) video from an image sequence of a dynamic scene captured while rapidly varying the exposure of each frame. Our approach consists of three parts: automatic exposure control during capture, HDR stitching across neighboring frames, and tonemapping for viewing. HDR stitching requires accurately registering neighboring frames and choosing appropriate pixels for computing the radiance map. We show examples for a variety of dynamic scenes. We also show how we can compensate for scene and camera movement when creating an HDR still from a series of bracketed still photographs.
Sing Bing Kang, Matthew Uyttendaele, Simon A. J. Winder, Richard Szeliski
ACM Trans. Graph.4
2002 On the Motion and Appearance of Specularities in Image Sequences
Rahul Swaminathan, Sing Bing Kang, Richard Szeliski, Antonio Criminisi, Shree K. Nayar
ECCV (1)3
2002 Symmetric Sub-Pixel Stereo Matching
Richard Szeliski, Daniel Scharstein
ECCV (2)1
2002 Modeling and Animating Realistic Faces from Images
Frédéric H. Pighin, Richard Szeliski, David Salesin
Int. J. Comput. Vis.2
2002 A Taxonomy and Evaluation of Dense Two-Frame Stereo Correspondence Algorithms
Daniel Scharstein, Richard Szeliski
Int. J. Comput. Vis.2
2002 Correction to Construction of Panoramic Image Mosaics with Global and Local Alignment
Harry Shum, Richard Szeliski
Int. J. Comput. Vis.2
2002 Video matting of complex scenes
abstract
This paper describes a new framework for video matting, the process of pulling a high-quality alpha matte and foreground from a video sequence. The framework builds upon techniques in natural image matting, optical flow computation, and background estimation. User interaction is comprised of garbage matte specification if background estimation is needed, and hand-drawn keyframe segmentations into "foreground," "background" and "unknown". The segmentations, called trimaps, are interpolated across the video volume using forward and backward optical flow. Competing flow estimates are combined based on information about where flow is likely to be accurate. A Bayesian matting technique uses the flowed trimaps to yield high-quality mattes of moving foreground elements with complex boundaries filmed by a moving camera. A novel technique for smoke matte extraction is also demonstrated.
Yung-Yu Chuang, Aseem Agarwala, Brian Curless, David Salesin, Richard Szeliski
ACM Trans. Graph.5
2001 A Bayesian Approach to Digital Matting
abstract
This paper proposes a new Bayesian framework for solving the matting problem, i.e. extracting a foreground element from a background image by estimating an opacity for each pixel of the foreground element. Our approach models both the foreground and background color distributions with spatially-varying sets of Gaussians, and assumes a fractional blending of the foreground and background colors to produce the final output. It then uses a maximum-likelihood criterion to estimate the optimal opacity, foreground and background simultaneously. In addition to providing a principled approach to the matting problem, our algorithm effectively handles objects with intricate boundaries, such as hair strands and fur, and provides an improvement over existing techniques for these difficult cases.
Yung-Yu Chuang, Brian Curless, David Salesin, Richard Szeliski
CVPR (2)4
2001 Handling Occlusions in Dense Multi-view Stereo
abstract
While stereo matching was originally formulated as the recovery of 3D shape from a pair of images, it is now generally recognized that using more than two images can dramatically improve the quality of the reconstruction. Unfortunately, as more images are added, the prevalence of semi-occluded regions (pixels visible in some but not all images) also increases. We propose some novel techniques to deal with this problem. Our first idea is to use a combination of shiftable windows and a dynamically selected subset of the neighboring images to do the matches. Our second idea is to explicitly label occluded pixels within a global energy minimization framework, and to reason about visibility within this framework so that only truly visible pixels are matched. Experimental results show a dramatic improvement using the first idea over conventional multibaseline stereo, especially when used in conjunction with a global energy minimization technique. These results also show that explicit occlusion labeling and visibility reasoning do help, but not significantly, if the spatial and temporal selection is applied first.
Sing Bing Kang, Richard Szeliski, Jinxiang Chai
CVPR (1)2
2001 Eliminating Ghosting and Exposure Artifacts in Image Mosaics
abstract
As panoramic photography becomes increasingly popular, there is a greater need for high-quality software to automatically create panoramic images. Existing algorithms either produce a rough "stitch" that cannot deal with common artifacts, or require user input. This paper presents methods for dealing with two artifacts that often occur in practice. Our first contribution is a method for dealing with objects that move between different views of a dynamic scene. If such moving objects are left in, they will appear blurry and "ghosted". Treating such regions as nodes in a graph, we use a vertex cover algorithm to selectively remove all but one instance of each object. Our second contribution is a method for continuously adjusting exposure across multiple images in order to eliminate visible shifts in brightness or hue. We compute exposure corrections on a block-by block basis, then smoothly interpolate the parameters using a spline to get spatially continuous exposure adjustment. Our enhancements, combined with previously published techniques for automatic image stitching, result in a high-quality automated stitcher that exhibits far fewer artifacts than existing software.
Matthew Uyttendaele, Ashley Eden, Richard Szeliski
CVPR (2)3
2001 Optimal Texture Map Reconstruction from Multiple Views
abstract
The recovery of 3D models from multiple reference images involves not only the extraction of 3D shape, but also of texture. Assuming that all surfaces are Lambertian, the resulting final texture is typically computed as a linear combination of reference textures. This is, however, not the optimal means for reconstructing textures, since this does not model the anisotropy in the texture projection. Furthermore, the spatial image sampling may be quite variable within a fore-shortened surface. This also has important implications for computer vision techniques that involve analysis by synthesis and the image-based rendering (IBR) technique of view-dependent texture mapping (VDTM). Starting with sampling theory, we show how weights should be spatially distributed for optimal texture construction. The local weights take into consideration the effects of anisotropy and variable spatial image sampling. We also present experimental results to verify our analysis.
Lifeng Wang 0001, Sing Bing Kang, Richard Szeliski, Harry Shum
CVPR (1)3
2001 An Integrated Bayesian Approach to Layer Extraction from Image Sequences
abstract
This paper describes a Bayesian approach for modeling 3D scenes as collection of approximately planar layers that are arbitrarily positioned and oriented in the scene. In contrast to much of the previous work on layer-based motion modeling, which computes layered descriptions of 2D image motion, our work leads to a 3D description of the scene. There are two contributions within the paper. The first is to formulate the prior assumptions about the layers and scene within a Bayesian decision making framework which is used to automatically determine the number of layers and the assignment of individual pixels to layers. The second is algorithmic. In order to achieve the optimization, a Bayesian version of RANSAC is developed with which to initialize the segmentation. Then, a generalized expectation maximization method is used to find the MAP solution.
Philip Torr 0001, Richard Szeliski, P. Anandan 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2000 Layer Extraction from Multiple Images Containing Reflections and Transparency
abstract
Many natural images contain reflections and transparency, i.e., they contain mixtures of reflected and transmitted light. When viewed from a moving camera, these appear as the superposition of component layer images moving relative to each other. The problem of multiple motion recovery has been previously studied by a number of researchers. However no one has yet demonstrated how to accurately recover the component images themselves. In this paper we develop an optimal approach to recovering layer images and their associated motions from an arbitrary number of composite images. We develop two different techniques for estimating the component layer images given known motion estimates. The first approach uses constrained least squares to recover the layer images. The second approach iteratively refines lower and upper bounds on the layer images using two novel compositing operations, namely minimum- and maximum-composites of aligned images. We combine these layer extraction techniques with a dominant motion estimator and a subsequent motion refinement stage. This results in a completely automated system that recovers transparent images and motions from a collection of input images.
Richard Szeliski, Shai Avidan, P. Anandan 0001
CVPR1
2000 The Geometry-Image Representation Tradeoff for Rendering
abstract
It is generally recognized that 3-D models are compact representations for rendering. While pure image-based rendering techniques are capable of producing highly photorealistic outputs, the size of the input "model" is usually very large. The important issues in trading off geometry versus images include compactness of representation, photorealism of reconstructed views, and speed of rendering. We describe our past work in modeling and rendering, and articulate lessons learnt. We then delineate our vision of an ideal rendering system.
Sing Bing Kang, Richard Szeliski, P. Anandan 0001
ICIP2
2000 Scene Reconstruction from Multiple Cameras
abstract
This paper reviews a number of previously developed stereo matching algorithms and representations. It focuses on techniques that are especially well suited for stereoscopic and 3-D imaging applications such as novel view generation and the mixing of live imagery with synthetic computer graphics. The paper reviews some approaches to the classic problem of recovering a depth map from two or more images. It then describes a number of newer representations (and their associated reconstruction algorithms), including volumetric representations, layered plane-plus-parallax representations, and multiple depth maps.
Richard Szeliski
ICIP1
2000 Environment matting extensions: towards higher accuracy and real-time capture
abstract
Environment matting is a generalization of traditional bluescreen matting. By photographing an object in front of a sequence of structured light backdrops, a set of approximate light-transport paths through the object can be computed. The original environment matting research chose a middle ground—using a moderate number of photographs to produce results that were reasonably accurate for many objects. In this work, we extend the technique in two opposite directions: recovering a more accurate model at the expense of using additional structured light backdrops, and obtaining a simplified matte using just a single backdrop. The first extension allows for the capture of complex and subtle interactions of light with objects, while the second allows for video capture of colorless objects in motion.
Yung-Yu Chuang, Douglas E. Zongker, Joel Hindorff, Brian Curless, David Salesin, Richard Szeliski
SIGGRAPH6
2000 Video textures
abstract
This paper introduces a new type of medium, called a video texture, which has qualities somewhere between those of a photograph and a video. A video texture provides a continuous infinitely varying stream of images. While the individual frames of a video texture may be repeated from time to time, the video sequence as a whole is never repeated exactly. Video textures can be used in place of digital photos to infuse a static image with dynamic qualities and explicit actions. We present techniques for analyzing a video clip to extract its structure, and for synthesizing a new, similar looking video of arbitrary length. We combine video textures with view morphing techniques to obtain 3D video textures. We also introduce video-based animation, in which the synthesis of video textures can be guided by a user through high-level interactive controls. Applications of video textures and their extensions include the display of dynamic scenes on web pages, the creation of dynamic backdrops for special effects and games, and the interactive control of video-based animation.
Arno Schödl, Richard Szeliski, David Salesin, Irfan A. Essa
SIGGRAPH2
2000 Systems and Experiment Paper: Construction of Panoramic Image Mosaics with Global and Local Alignment
Harry Shum, Richard Szeliski
Int. J. Comput. Vis.2
1999 Stereo Algorithms and Representations for Image-based Rendering
abstract
This paper reviews a number of recently developed stereo matching algorithms and representations. It focuses on techniques that are especially well suited for image-based rendering applications such as novel view generation and the mixing of live imagery with synthetic computer graphics. The paper begins by reviewing some recent approaches to the classic problem of recovering a depth map from two or more images. It then describes a number of newer representations (and their associated reconstruction algorithms), including volumetric representations, layered plane-plus-parallax representations, and multiple depth maps. Each of these techniques has its own strengths and weaknesses, which are discussed.
Richard Szeliski
BMVC1
1999 A Multi-View Approach to Motion and Stereo
abstract
This paper presents a new approach to computing dense depth and motion estimates from multiple images. Rather than computing a single depth or motion map from such a collection, we associate motion or depth estimates with each image in the collection (or at least some subset of the images). This has the advantage that the depth or motion of regions occluded in one image will still be represented in some other image. Thus, tasks such as novel view interpolation or motion-compensated prediction can be solved with greater fidelity. Furthermore, the natural variation in appearance between different images can be captured. To formulate motion and structure recovery, we cast the problem as a global optimization over the unknown motion or depth maps, and use robust smoothness constraints to constrain the space of possible solutions. We develop and evaluate some motion and depth estimation algorithms based on this framework.
Richard Szeliski
CVPR1
1999 Resynthesizing Facial Animation through 3D Model-based Tracking
abstract
Given video footage of a person's face, we present new techniques to automatically recover the face position and the facial expression from each frame in the video sequence. A 3D face model is fitted to each frame using a continuous optimization technique. Our model is based on a set of 3D face models that are linearly combined using 3D morphing. Our method has the advantages over previous techniques of fitting directly a realistic 3-dimensional face model and of recovering parameters that can be used directly in an animation system. We also explore many applications, including performance-driven animation (applying the recovered position and expression of the face to a synthetic character to produce an animation that mimics the input video), relighting the face, varying the camera position, and adding facial ornaments such as tattoos and scars.
Frédéric H. Pighin, Richard Szeliski, David Salesin
ICCV2
1999 Stereo Reconstruction from Multiperspective Panoramas
abstract
The paper presents a new approach to computing depth maps from a large collection of images where the camera motion has been constrained to planar concentric circles. We resample the resulting collection of regular perspective images into a set of multiperspective panoramas, and then compute depth maps directly from these resampled images. Only a small number of multiperspective panoramas is needed to obtain a dense and accurate 3D reconstruction, since our panoramas sample uniformly in three dimensions: rotation angle, inverse radial distance, and vertical elevation. Using multiperspective panoramas avoids the limited overlap between the original input images that causes problems in conventional multi-baseline stereo. Our approach differs from stereo matching of panoramic images taken from different locations, where the epipolar constraints are sine curves. For our multiperspective panoramas, the epipolar geometry to first order consists of horizontal lines. Therefore, any traditional stereo algorithm can be applied to multiperspective panoramas without modification. Experimental results show that our approach generates good depth maps that can be used for image based rendering tasks such as view interpolation and extrapolation.
Harry Shum, Richard Szeliski
ICCV2
1999 Prediction Error as a Quality Metric for Motion and Stereo
abstract
This paper presents a new methodology for evaluating the quality of motion estimation and stereo correspondence algorithms. Motivated by applications such as novel view generation and motion-compensated compression, we suggest that the ability to predict new views or frames is a natural metric for evaluating such algorithms. Our new metric has several advantages over comparing algorithm outputs to true motions or depths. First of all, it does not require the knowledge of ground truth data, which may be difficult or laborious to obtain. Second, it more closely matches the ultimate requirements of the application, which are typically tolerant of errors in uniform color regions, but very sensitive to isolated pixel errors or disocclusion errors. In the paper we develop a number of error metrics based on this paradigm, including forward and inverse prediction errors, residual motion error and local motion-compensated prediction error. We show results on a number of widely used motion and stereo sequences, many of which do not have associated ground truth data.
Richard Szeliski
ICCV1
1999 An Integrated Bayesian Approach to Layer Extraction from Image Sequences
abstract
This paper describes a Bayesian approach for modeling 3D scenes as a collection of approximately planar layers that are arbitrarily positioned and oriented in the scene. In contrast to much of the previous work on layer based motion modeling, which compute layered descriptions of 2D image motion, our work leads to a 3D description of the scene. We focus on the key problem of automatically segmenting the scene into layers as a precursor to recovery of stereo disparity data. The prior assumptions about the scene are formulated within a Bayesian decision making framework, and are then used to automatically determine the number of layers and the assignment of individual pixels to layers. Although using a collection of 3D layers has been previously proposed as an efficient and effective representation for multimedia applications, results to date have relied on hand segmentation. In contrast, the work described aims at fully automatic segmentation.
Philip Torr 0001, Richard Szeliski, P. Anandan 0001
ICCV2
1999 The Videomouse: A Camera-Based Multi-degree-of-freedom Input Device
abstract
The VideoMouse is a mouse that uses a camera as its input sensor. A real-time vision algorithm determines the six degree-of-freedom mouse posture, consisting of 2D motion, tilt in the forward/back and left/right axes, rotation of the mouse about its vertical axis, and some limited height sensing. Thus, a familiar 2D device can be extended for three-dimensional manipulation, while remaining suitable for standard 2D GUI tasks. We describe techniques for mouse functionality, 3D manipulation, navigating large 2D spaces, and using the camera for lightweight scanning tasks.
Ken Hinckley, Mike Sinclair, Erik Hanson, Richard Szeliski, Matthew Conway
ACM Symposium on User Interface Software and Technology4
1999 Stereo Matching with Transparency and Matting
Richard Szeliski, Polina Golland
Int. J. Comput. Vis.1
1998 A Layered Approach to Stereo Reconstruction
abstract
We propose a framework for extracting structure from stereo which represents the scene as a collection of approximately planar layers. Each layer consists of an explicit 3D plane equation, a colored image with per-pixel opacity (a sprite), and a per-pixel depth offset relative to the plane. Initial estimates of the layers are recovered using techniques taken from parametric motion estimation. These initial estimates are then refined using a re-synthesis algorithm which takes into account both occlusions and mixed pixels. Reasoning about such effects allows the recovery of depth and color information with high accuracy even in partially occluded regions. Another important benefit of our framework is that the output consists of a collection of approximately planar regions, a representation which is far more appropriate than a dense depth map for many applications such as rendering and video parsing.
Simon Baker, Richard Szeliski, P. Anandan 0001
CVPR2
1998 Interactive Construction of 3D Models from Panoramic Mosaics
abstract
This paper presents an interactive modeling system that constructs 3D models from a collection of panoramic image mosaics. A panoramic mosaic consists of a set of images taken around the same viewpoint, and a transformation matrix associated with each input image. Our system first recovers the camera pose for each mosaic from known line directions and points, and then constructs the 3D model using all available geometrical constraints. We partition constraints into soft and hard linear constraints so that the modeling process can be formulated as a linearly-constrained least-squares problem, which can be solved efficiently using QR factorization. The results of extracting wire frame and texture-mapped 3D models from single and multiple panoramas are presented.
Harry Shum, Richard Szeliski
CVPR3
1998 Construction and Refinement of Panoramic Mosaics with Global and Local Alignment
abstract
This paper presents techniques for constructing full view panoramic mosaics form sequences of images. Our representation associates a rotation matrix (and optionally a focal length) with each input image, rather than explicitly projecting all of the images onto a common surface (e.g., a cylinder). In order to reduce accumulated registration errors we apply global alignment (block adjustment) to whole sequence of images, which results in an optimal image mosaic (in the least squares sense). To compensate for small amounts of motion parallax introduced by translations of the camera and other unmodeled distortions we develop a local alignment (deghosting) technique which warps each image based on the results of pairwise local image registrations. By combining both global and local alignment we significantly improve the quality of our image mosaics thereby enabling the creation of full view panoramic mosaics with hand-held cameras.
Harry Shum, Richard Szeliski
ICCV2
1998 Stereo Matching with Transparency and Matting
abstract
This paper formulates and solves a new variant of the stereo correspondence problem: simultaneously recovering the disparities, true colors, and opacities of visible surface elements. This problem arises in newer applications of stereo reconstruction, such as view interpolation and the layering of real imagery with synthetic graphics for special effects and virtual studio applications. While this problem is intrinsically more difficult than traditional stereo correspondence, where only the disparities are being recovered, it provides a principled way of dealing with commonly occurring problems such as occlusions and the handling of mixed (foreground/background) pixels near depth discontinuities. It also provides a novel means for separating foreground and background objects (matting), without the use of a special blue screen. We formulate the problem as the recovery of colors and opacities in a generalized 3-D (x, y, d) disparity space, and solve the problem using a combination of initial evidence aggregation followed by iterative energy minimization.
Richard Szeliski, Polina Golland
ICCV1
1998 Synthesizing Realistic Facial Expressions from Photographs
abstract
Article Free Access Share on Synthesizing realistic facial expressions from photographs Authors: Frédéric Pighin Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile , Jamie Hecker Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile , Dani Lischinski Hebrew Univ. Hebrew Univ.View Profile , Richard Szeliski Microsoft Research Microsoft ResearchView Profile , David H. Salesin Univ. of Washington, Seattle Univ. of Washington, SeattleView Profile Authors Info & Claims SIGGRAPH '98: Proceedings of the 25th annual conference on Computer graphics and interactive techniquesJuly 1998 Pages 75–84https://doi.org/10.1145/280814.280825Published:24 July 1998Publication History 446citation2,172DownloadsMetricsTotal Citations446Total Downloads2,172Last 12 Months71Last 6 weeks7 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my AlertsNew Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteeReaderPDF
Frédéric H. Pighin, Jamie Hecker, Dani Lischinski, Richard Szeliski, David Salesin
SIGGRAPH4
1998 Layered Depth Images
abstract
In this paper we present a set of efficient image based rendering methods capable of rendering multiple frames per second on a PC. The first method warps Sprites with Depth representing smooth surfaces without the gaps found in other techniques. A second method for more general scenes performs warping from an intermediate representation called a Layered Depth Image (LDI). An LDI is a view of the scene from a single input camera view, but with multiple pixels along each line of sight. The size of the representation grows only linearly with the observed depth complexity in the scene. Moreover, because the LDI data are represented in a single image coordinate system, McMillan's warp ordering algorithm can be successfully adapted. As a result, pixels are drawn in the output image in back-to-front order. No z-buffer is required, so alphacompositing can be done efficiently without depth sorting. This makes splatting an efficient solution to the resampling problem. 1 Introduction Image base...
Jonathan Shade, Steven J. Gortler, Li-wei He, Richard Szeliski
SIGGRAPH4
1998 Interactive 3D modeling from multiple images using scene regularities
abstract
Due to the complexity of real scenes and the fragility of fully automated vision techniques, results from many automated modeling systems are disappointing. Automated techniques often require manual clean-up and postprocessing to segment the scene into coherent objects and surfaces, or to triangulate sparse point matches. They may also be required to enforce geometric constraints such as known orientations of surfaces. For instance, building interiors and exteriors provide vertical and horizontal lines and parallel and perpendicular planes. In this paper, we attack the 3D modeling problem from the other side: we specify some geometric knowledge ahead of time (e.g., known orientations of lines, co-planarity of points, initial scene segmentations), and use these constraints to guide our matching and reconstruction algorithms. We present two interactive (semi-automated) systems for recovering 3D models of large-scale environments from multiple images.
Harry Shum, Richard Szeliski, Simon Baker, P. Anandan 0001
WACV2
1998 Stereo Matching with Nonlinear Diffusion
Daniel Scharstein, Richard Szeliski
Int. J. Comput. Vis.2
1998 Robust Shape Recovery from Occluding Contours Using a Linear Smoother
Richard Szeliski, Richard Weiss 0001
Int. J. Comput. Vis.1
1997 Creating full view panoramic image mosaics and environment maps
abstract
This paper presents a novel approach to creating full view panoramic mosaics from image sequences.Unlike current panoramic stitching methods, which usually require pure horizontal camera panning, our system does not require any controlled motions or constraints on how the images are taken (as long as there is no strong motion parallax).For example, images taken from a hand-held digital camera can be stitched seamlessly into panoramic mosaics.Because we represent our image mosaics using a set of transforms, there are no singularity problems such as those existing at the top and bottom of cylindrical or spherical maps.Our algorithm is fast and robust because it directly recovers 3D rotations instead of general 8 parameter planar perspective transforms.Methods to recover camera focal length are also presented.We also present an algorithm for efficiently extracting environment maps from our image mosaics.By mapping the mosaic onto an artibrary texture-mapped polyhedron surrounding the origin, we can explore the virtual environment using standard 3D graphics viewers and hardware without requiring special-purpose players.
Richard Szeliski, Harry Shum
SIGGRAPH1
1997 A Parallel Feature Tracker for Extended Image Sequences
Sing Bing Kang, Richard Szeliski, Harry Shum
Comput. Vis. Image Underst.2
1997 3-D Scene Data Recovery Using Omnidirectional Multibaseline Stereo
Sing Bing Kang, Richard Szeliski
Int. J. Comput. Vis.2
1997 Spline-Based Image Registration
Richard Szeliski, James M. Coughlan
Int. J. Comput. Vis.1
1997 Shape Ambiguities in Structure From Motion
abstract
This paper examines the fundamental ambiguities and uncertainties inherent in recovering structure from motion. By examining the eigenvectors associated with null or small eigenvalues of the Hessian matrix, we can quantify the exact nature of these ambiguities and predict how they affect the accuracy of the reconstructed shape. Our results for orthographic cameras show that the bas-relief ambiguity is significant even with many images, unless a large amount of rotation is present. Similar results for perspective cameras suggest that three or more frames and a large amount of rotation are required for metrically accurate reconstruction.
Richard Szeliski, Sing Bing Kang
IEEE Trans. Pattern Anal. Mach. Intell.1
1997 A layered video object coding system using sprite and affine motion model
abstract
A layered video object coding system is presented in this paper. The goal is to improve video coding efficiency by exploiting the layering of video and to support content-based functionality. These two objectives are accomplished using a sprite technique and an affine motion model on a per-object basis. Several novel algorithms have been developed for mask processing and coding, trajectory coding, sprite accretion and coding, locally affine motion compensation, error signal suppression, and image padding. Compared with conventional frame-based coding methods, better experimental results on both hybrid and natural scenes have been obtained using our coding scheme. We also demonstrate content-based functionality which can be easily achieved in our system.
Ming-Chieh Lee, Wei-Ge Chen, Chih-lung Bruce Lin, Chuang Gu, Tomislav Markoc, Steven I. Zabinsky, Richard Szeliski
IEEE Trans. Circuits Syst. Video Technol.7
1996 3-D Scene Data Recovery using Omnidirectional Multibaseline Stereo
abstract
A traditional approach to extracting geometric information from a large scene is to compute multiple 3-D depth maps from stereo pairs or direct range finders, and then to merge the 3-D data. However, the resulting merged depth maps may be subject to merging errors if the relative poses between depth maps are not known exactly. In addition, the 3-D data may also have to be resampled before merging, which adds additional complexity and potential sources of errors. This paper provides a means of directly extracting 3-D data covering a very wide field of view, thus by-passing the need for numerous depth map merging. In our work, cylindrical images are first composited from sequences of images taken while the camera is rotated 360/spl deg/ about a vertical axis. By taking such image panoramas at different camera locations, we can recover 3-D data of the scene using a set of simple techniques: feature tracking, an 8-point structure from motion algorithm, and multibaseline stereo. We also investigate the effect of median filtering on the recovered 3-D point distributions, and show the results of our approach applied to both synthetic and real scenes.
Sing Bing Kang, Richard Szeliski
CVPR2
1996 Stereo Matching with Non-Linear Diffusion
abstract
One of the central problems in stereo matching (and other image registration tasks) is the selection of optimal window sizes for comparing image regions. This paper addresses this problem with some novel algorithms based on iteratively diffusing support at different disparity hypotheses, and locally controlling the amount of diffusion based on the current quality of the disparity estimate. It also develops a novel Bayesian estimation technique which significantly outperforms techniques based on area-based matching (SSD) and regular diffusion. We provide experimental results on both synthetic and real stereo image pairs.
Daniel Scharstein, Richard Szeliski
CVPR2
1996 Shape Ambiguities in Structure from Motion
Richard Szeliski, Sing Bing Kang
ECCV (1)1
1996 The Lumigraph
abstract
This paper discusses a new method for capturing the complete appearance of bothsynthetic and real world objects and scenes,representing this information, and then using this representation to render images of the object from new camera positions.Unlike the shape capture process traditionally used in computer vision and the rendering process traditionally used in computer graphics, our approach does not rely on geometric representations.Instead we sample and reconstruct a 4D function.which we call a Lumigraph.The Lumigraph is a subset o f the complete plenoptic fu nction that describes the flow of light at all positions in all directions.With the Lumigraph.new images of the object can be generated very quick]y,independentof the geometric or illumination complexity of the scene or object.The paper discusses a complete working system including the capture of sainples.the construction of the Lumigraph, and thesubsequent rendering of images from this new representation.even i f given accurate geometric models.Quicktime V R [6] was one ofthe first systems to suggest that the traditional modeling/rendering process can beskipped.Instead.
Steven J. Gortler, Radek Grzeszczuk, Richard Szeliski, Michael F. Cohen
SIGGRAPH3
1996 Matching 3-D anatomical surfaces with non-rigid deformations using octree-splines
Richard Szeliski, Stéphane Lavallée
Int. J. Comput. Vis.1
1996 Motion Estimation with Quadtree Splines
abstract
This paper presents a motion estimation algorithm based on a new multiresolution representation, the quadtree spline. This representation describes the motion field as a collection of smoothly connected patches of varying size, where the patch size is automatically adapted to the complexity of the underlying motion. The topology of the patches is determined by a quadtree data structure, and both split and merge techniques are developed for estimating this spatial subdivision. The quadtree spline is implemented using another novel representation, the adaptive hierarchical basis spline, and combines the advantages of adaptively-sized correlation windows with the speedups obtained with hierarchical basis preconditioners. Results are presented on some standard motion sequences.
Richard Szeliski, Harry Shum
IEEE Trans. Pattern Anal. Mach. Intell.1
1995 Motion Estimation with Quadtree Splines
abstract
This paper presents a motion estimation algorithm based on a new multiresolution representation, the quadtree spline. This representation describes the motion field as a collection of smoothly connected patches of varying size, where the patch size is automatically adapted to the complexity of the underlying motion. The topology of the patches is determined by a quadtree data structure, and both split and merge techniques are developed for estimating this spatial subdivision. The quadtree spline is implemented using another novel representation, the adaptive hierarchical basis spline, and combines the advantages of adaptively-sized correlation windows with the speedups obtained with hierarchical basis preconditioners. Results are presented on some standard motion sequences.>
Richard Szeliski, Harry Shum
ICCV1
1995 Recovering the Position and Orientation of Free-Form Objects from Image Contours Using 3D Distance Maps
abstract
The accurate matching of 3D anatomical surfaces with sensory data such as 2D X-ray projections is a basic problem in computer and robot assisted surgery, In model-based vision, this problem can be formulated as the estimation of the spatial pose (position and orientation) of a 3D smooth object from 2D video images. The authors present a new method for determining the rigid body transformation that describes this match. The authors' method performs a least squares minimization of the energy necessary to bring the set of the camera-contour projection lines tangent to the surface. To correctly deal with projection lines that penetrate the surface, the authors consider the minimum signed distance to the surface along each line (i.e., distances inside the object are negative). To quickly and accurately compute distances to the surface, the authors introduce a precomputed distance map represented using an octree spline whose resolution increases near the surface. This octree structure allows the authors to quickly find the minimum distance along each line using best-first search. Experimental results for 3D surface to 2D projection matching are presented for both simulated and real data. The combination of the authors' problem formulation in 3D, their computation of line to surface distances with the octree-spline distance map, and their simple minimization technique based on the Levenberg-Marquardt algorithm results in a method that solves the 3D/2D matching problem for arbitrary smooth shapes accurately and quickly.>
Stéphane Lavallée, Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.2
1994 Hierarchical spline-based image registration
abstract
The problem of image registration subsumes a number of topics in multiframe image analysis, including the computation of optic flow (general pixel-based motion), stereo correspondence, structure from motion, and feature tracking. We present a new registration algorithm based on a spline representation of the displacement field which can be specialized to solve all of the above mentioned problems. In particular, we show how to compute local flow, global (parametric) flow, rigid flow resulting from camera egomotion, and multiframe versions of the above problems. Using a spline-based description of the flow removes the need for overlapping correlation windows, and produces an explicit measure of the correlation between adjacent flow estimates. We demonstrate our algorithm on multiframe image registration and the recovery of 3D projective scene geometry. We also provide results on a number of standard motion sequences.>
Richard Szeliski, James M. Coughlan
CVPR1
1994 Image mosaicing for tele-reality applications
abstract
This paper presents some techniques for automatically deriving realistic 2-D scenes and 3-D geometric models from video sequences. These techniques can be used to build environments and 3-D models for virtual reality application based on recreating a true scene, i.e., tele-reality applications. The fundamental technique used in this paper is image mosaicing, i.e., the automatic alignment of multiple images into larger aggregates which are then used to represent portions of a 3-D scene. The paper first examines the easiest problems, those of flat scene and panoramic scene mosaicing. It then progresses to more complicated scenes with depth, and concludes with full 3-D models. The paper also discusses a number of novel applications based on tele-reality technology.>
Richard Szeliski
WACV1
1994 Recovering 3D Shape and Motion from Image Streams Using Nonlinear Least Squares
Richard Szeliski, Sing Bing Kang
J. Vis. Commun. Image Represent.1
1993 Recovering 3D shape and motion from image streams using nonlinear least squares
abstract
A shape and motion estimation algorithm based on nonlinear least squares applied to the tracks of features through time is presented. While the authors' approach requires iteration, it quickly converges to the desired solution, even in the absence of a priori knowledge about the shape or motion. Important features of the algorithm include its ability to handle partial point tracks and true perspective, its ability to use line segment matches and point matches simultaneously, and its use of an object-centered representation for faster and more accurate structure and motion recovery.>
Richard Szeliski, Sing Bing Kang
CVPR1
1993 Modeling surfaces of arbitrary topology with dynamic particles
abstract
A new approach to surface modeling and reconstruction is developed which overcomes some important limitations of existing surface representations methods. The approach features two components. The first is a dynamic self-organizing oriented particle system which discovers topological and geometric surface structure implicit in visual data. The oriented particles evolve according to Newtonian mechanics and interact through long-range attraction forces, short-range repulsion forces, and coplanarity, conormality, and cocircularity forces. The second component is an efficient triangulation scheme that connects the particles into a continuous global surface model that is consistent with the inferred structure. A flexible surface reconstruction algorithm is developed that can compute complete, detailed, viewpoint-invariant geometric surface descriptions of objects with arbitrary topology. The algorithms are applied to 3-D medical image segmentation and to surface reconstruction from object silhouettes.>
Richard Szeliski, David Tonnesen, Demetri Terzopoulos
CVPR1
1993 Robust shape recovery from occluding contours using a linear smoother
abstract
Recovering the shape of an object from triangulation fails at occluding contours of smooth objects because the contour generators are view dependent. For three or more views, shape recovery is possible, and several algorithms have been developed for this purpose. The authors' approach uses a linear smoother to optimally combine all of the measurements available at the contours (and other edges) in all of the images. This allows extraction of a robust and dense estimate of surface shape, and shape information from both surface markings and occluding contours.>
Richard Szeliski, Richard Weiss 0001
CVPR1
1993 Impossible Shaded Images
abstract
It is shown that shaded images that cannot have originated from a uniformly illuminated, smooth continuous surface with uniform albedo exist. The typical condition where this occurs is when a dark area (corresponding to a region of high gradient) is surrounded by a lighter region (with low gradient). For this to correspond to a real surface, it must be established that there is a local extremum or area of lower gradient inside the dark region. This, in turn, will show up as either a light area in the image or an orientation discontinuity in the surface (thus violating either intensity or smoothness constraints). The impossibility of a shaded image can be established by counting the number of extrema inside a region corresponding to an isolated surface patch.>
Berthold K. P. Horn, Richard Szeliski, Alan L. Yuille
IEEE Trans. Pattern Anal. Mach. Intell.2
1992 From accurate range imaging sensor calibration to accurate model-based 3D object localization
abstract
The registration of multiple 3D data sets obtained with a laser range finder is examined. A sensor calibration technique based on the conjunction of a mathematical camera model, the N-planes B-spline (NPBS), with an accurate mechanical calibration setup, is proposed. An algorithm for recovering the rigid transformation (rotation and translation) between the two sets of 3D coordinates obtained by the calibrated imaging range sensor is developed. Its input is sets of points lying on the surface of the object. This input is converted into an octree-spline that allows point-to-surface distances to be computed quickly. A robust nonlinear least-squares minimization technique then finds the optimal pose by minimizing the sum of square distances between the two sets of 3D coordinates. The algorithm has been applied to matching human faces with highly accurate results.>
Guillaume Champleboux, Stéphane Lavallée, Richard Szeliski, Lionel Brunie
CVPR3
1992 Using Force Fields Derived from 3D Distance Maps for Inferring the Attitude of a 3D Rigid Object
Lionel Brunie, Stéphane Lavallée, Richard Szeliski
ECCV3
1992 Surface modeling with oriented particle systems
abstract
Splines and deformable surface models are widely used in computer graphics to describe free-form surfaces. These methods require manual preprocessing to discretize the surface into patches and to specify their connectivity. We present a new model of elastic surfaces based on interacting particle systems, which, unlike previous techniques, can be used to sptiL join, or extend surfaces without the need for manual intervention. The particles we use have longrange attraction forces and short-range repulsion forces and follow Newtonian dynamics, much tiie recent computational models of fluids and solids. To enable our particles to model surface elements instead of point masses or volume elements, we add an orientation to each particle’s state. We devise new interaction potentials for our oriented particles which favor locally planar or spherical arrangements. We also develop techniques for adding new particles automatically, which enables our surfaces to stretch and grow. We demonstrate the application of our new particle system to modeling surfaces in 3-D and the interpolation of 3-D point sets.
Richard Szeliski, David Tonnesen
SIGGRAPH1
1991 Shape from rotation
abstract
The construction of a 3D surface model of an object rotating in front of a camera is examined. Previous research in depth from motion has demonstrated the power of using an incremental approach to depth estimation. The author extends this approach to more general motion and uses a full 3D surface model instead of a 2 1/2 D depth map. The algorithm starts with a flow field computed using local correlation. It then projects individual measurements into 3D points with associated uncertainties. Nearby points from successive frames are merged to improve the position estimates. These points are then used to construct a deformable surface model, which is refined over time. The application of novel techniques to several image sequences is demonstrated.>
Richard Szeliski
CVPR1
1991 Fast shape from shading
Richard Szeliski
CVGIP Image Underst.1
1990 Fast Shape from Shading
Richard Szeliski
ECCV1
1990 Bayesian modeling of uncertainty in low-level vision
Richard Szeliski
Int. J. Comput. Vis.1
1990 Fast Surface Interpolation Using Hierarchical Basis Functions
abstract
An alternative to multigrid relaxation that is much easier to implement and more generally applicable is presented. Conjugate gradient descent is used in conjunction with a hierarchical (multiresolution) set of basis functions. The resultant algorithm uses a pyramid to smooth the residual vector before the direction is computed. Simulation results showing the speed of convergence and its dependence on the choice of interpolator, the number of smoothing levels, and other factors are presented. The relationship of this approach to other multiresolution relaxation and representation schemes is also discussed.>
Richard Szeliski
IEEE Trans. Pattern Anal. Mach. Intell.1
1989 Fast surface interpolation using hierarchical basis functions
abstract
The rapid solution of surface interpolation and other regularization problems on massively parallel architectures is an important problem within computer vision. Fast relaxation algorithms can be used to integrate sparse data, resolve ambiguities in optic flow fields, and guide stereo matching algorithms. In the present paper, an alternative to multigrid relaxation which is much easier to implement is presented. This approach uses conjugate-gradient descent in conjunction with a hierarchical (multiresolution) set of basis functions. The resulting algorithm uses a pyramid to smooth the residual vector before the new direction is computed. Simulation results show the speed and its dependence on the choice of interpolator, the number of smoothing levels, and other factors. Also discussed is relationship of this approach to other multiresolution relaxation and representation schemes.>
Richard Szeliski
CVPR1
1989 From splines to fractals
abstract
Deterministic splines and stochastic fractals are complementary techniques for generating free-form shapes. Splines are easily constrained and well suited to modeling smooth, man-made objects. Fractals, while difficult to constrain, are suitable for generating various irregular shapes found in nature. This paper develops constrained fractals, a hybrid of splines and fractals which intimately combines their complementary features. This novel shape synthesis technique stems from a formal connection between fractals and generalized energy-minimizing splines which may be derived through Fourier analysis. A physical interpretation of constrained fractal generation is to drive a spline subject to constraints with modulated white noise, letting the spline diffuse the noise into the desired fractal spectrum as it settles into equilibrium. We use constrained fractals to synthesize realistic terrain models from sparse elevation data.
Richard Szeliski, Demetri Terzopoulos
SIGGRAPH1
1989 Kalman filter-based algorithms for estimating depth from image sequences
Larry H. Matthies, Takeo Kanade, Richard Szeliski
Int. J. Comput. Vis.3
1989 An Analysis of the Elastic Net Approach to the Traveling Salesman Problem
abstract
This paper analyzes the elastic net approach (Durbin and Willshaw 1987) to the traveling salesman problem of finding the shortest path through a set of cities. The elastic net approach jointly minimizes the length of an arbitrary path in the plane and the distance between the path points and the cities. The tradeoff between these two requirements is controlled by a scale parameter K. A global minimum is found for large K, and is then tracked to a small value. In this paper, we show that (1) in the small K limit the elastic path passes arbitrarily close to all the cities, but that only one path point is attracted to each city, (2) in the large K limit the net lies at the center of the set of cities, and (3) at a critical value of K the energy function bifurcates. We also show that this method can be interpreted in terms of extremizing a probability distribution controlled by K. The minimum at a given K corresponds to the maximum a posteriori (MAP) Bayesian estimate of the tour under a natural statistical interpretation. The analysis presented in this paper gives us a better understanding of the behavior of the elastic net, allows us to better choose the parameters for the optimization, and suggests how to extend the underlying ideas to other domains.
Richard Durbin, Richard Szeliski, Alan L. Yuille
Neural Comput.2
1988 Incremental estimation of dense depth maps from image sequences
abstract
The authors introduce a novel pixel-based (iconic) algorithm that estimates depth and depth uncertainty at each pixel and incrementally refines these estimates over time. They describe the algorithm for translations parallel to the image plane and contrast its formulation and performance to that of a feature-based Kalman filtering algorithm. They compare the performance of the two approaches by analyzing their theoretical convergence rates, by conducting quantitative experiments with images of a flat poster, and by conducting qualitative experiments with images of a realistic outdoor scene model. The results show that the method is an effective way to extract depth from lateral camera translations and suggest that it will play an important role in low-level vision.>
Larry H. Matthies, Richard Szeliski, Takeo Kanade
CVPR2
1988 Estimating Motion From Sparse Range Data Without Correspondence
Richard Szeliski
ICCV1
1987 Regularization Uses Fractal Priors
Richard Szeliski
AAAI1