Jan-Michael Frahm

dblp:19/6011 · DBLP profile ↗
← Back
114ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 91 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 78 · 4 first-author · 3 since 2021Security and privacy · 8 · 1 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-authorSystems, architecture and hardware · 3Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021
YearPublicationVenuePosition
2024 Supervision Interpolation via LossMix: Generalizing Mixup for Object Detection and Beyond
abstract
The success of data mixing augmentations in image classification tasks has been well-received. However, these techniques cannot be readily applied to object detection due to challenges such as spatial misalignment, foreground/background distinction, and plurality of instances. To tackle these issues, we first introduce a novel conceptual framework called Supervision Interpolation (SI), which offers a fresh perspective on interpolation-based augmentations by relaxing and generalizing Mixup. Based on SI, we propose LossMix, a simple yet versatile and effective regularization that enhances the performance and robustness of object detectors and more. Our key insight is that we can effectively regularize the training on mixed data by interpolating their loss errors instead of ground truth labels. Empirical results on the PASCAL VOC and MS COCO datasets demonstrate that LossMix can consistently outperform state-of-the-art methods widely adopted for detection. Furthermore, by jointly leveraging LossMix with unsupervised domain adaptation, we successfully improve existing approaches and set a new state of the art for cross-domain object detection.
Thanh Vu 0001, Baochen Sun, Bodi Yuan, Alex Ngai, Jan-Michael Frahm
AAAI6
2023 A Practical Stereo Depth System for Smart Glasses
abstract
We present the design of a productionized end-to-end stereo depth sensing system that does pre-processing, online stereo rectification, and stereo depth estimation with a fallback to monocular depth estimation when rectification is unreliable. The output of our depth sensing system is then used in a novel view generation pipeline to create 3D computational photography effects using point-of-view images captured by smart glasses. All these steps are executed on-device on the stringent compute budget of a mobile phone, and because we expect the users can use a wide range of smartphones, our design needs to be general and cannot be dependent on a particular hardware or ML accelerator such as a smartphone GPU. Although each of these steps is well studied, a description of a practical system is still lacking. For such a system, all these steps need to work in tandem with one another and fallback gracefully on failures within the system or less than ideal input data. We show how we handle unforeseen changes to calibration, e.g., due to heat, robustly support depth estimation in the wild, and still abide by the memory and latency constraints required for a smooth user experience. We show that our trained models are fast, and run in less than 1s on a six-year-old Samsung Galaxy S8 phone's CPU. Our models generalize well to unseen data and achieve good results on Middlebury and in-the-wild images captured from the smart glasses.
Jialiang Wang 0001, Daniel Scharstein, Akash Bapat, Kevin Matzen, Matthew Yu, Jonathan Lehman, Suhib Alsisan, Yanghan Wang, Sam S. Tsai, Jan-Michael Frahm, Peter Vajda, Michael F. Cohen, Matthew Uyttendaele
CVPR10
2023 MVPSNet: Fast Generalizable Multi-view Photometric Stereo
abstract
We propose a fast and generalizable solution to Multiview Photometric Stereo (MVPS), called MVPSNet. The key to our approach is a feature extraction network that effectively combines images from the same view captured under multiple lighting conditions to extract geometric features from shading cues for stereo matching. We demonstrate these features, termed ‘Light Aggregated Feature Maps’ (LAFM), are effective for feature matching even in textureless regions, where traditional multi-view stereo methods often fail. Our method produces similar reconstruction results to PS-NeRF, a state-of-the-art MVPS method that optimizes a neural network per-scene, while being 411× faster (105 seconds vs. 12 hours) in inference. Additionally, we introduce a new synthetic dataset for MVPS, sMVPS, which is shown to be effective for training a generalizable MVPS method.
Dongxu Zhao 0001, Daniel Lichy, Pierre-Nicolas Perrin, Jan-Michael Frahm, Roni Sengupta
ICCV4
2023 Toward Edge-Efficient Dense Predictions with Synergistic Multi-Task Neural Architecture Search
abstract
In this work, we propose a novel and scalable solution to address the challenges of developing efficient dense predictions on edge platforms. Our first key insight is that MultiTask Learning (MTL) and hardware-aware Neural Architecture Search (NAS) can work in synergy to greatly benefit on-device Dense Predictions (DP). Empirical results reveal that the joint learning of the two paradigms is surprisingly effective at improving DP accuracy, achieving superior performance over both the transfer learning of single-task NAS and prior state-of-the-art approaches in MTL, all with just 1/10th of the computation. To the best of our knowledge, our framework, named EDNAS, is the first to successfully leverage the synergistic relationship of NAS and MTL for DP. Our second key insight is that the standard depth training for multi-task DP can cause significant instability and noise to MTL evaluation. Instead, we propose JAReD, an improved, easy-to-adopt Joint Absolute-Relative Depth loss, that reduces up to 88% of the undesired noise while simultaneously boosting accuracy. We conduct extensive evaluations on standard datasets, benchmark against strong baselines and state-of-the-art approaches, as well as provide an analysis of the discovered optimal architectures.
Thanh Vu 0001, Yanqi Zhou, Chunfeng Wen, Jan-Michael Frahm
WACV5
2022 Leveraging Disentangled Representations to Improve Vision-Based Keystroke Inference Attacks Under Low Data Constraints
abstract
Keystroke inference attacks are a form of side-channel attacks in which an attacker leverages various techniques to recover a user's keystrokes as she inputs information into some display (e.g., while sending a text message or entering her pin). Typically, these attacks leverage machine learning approaches, but assessing the realism of the threat space has lagged behind the pace of machine learning advancements, due in-part, to the challenges in curating large real-life datasets. We aim to overcome the challenge of having limited number of real data by introducing a video domain adaptation technique that is able to leverage synthetic data through supervised disentangled learning. Specifically, for a given domain, we decompose the observed data into two factors of variation: Style and Content. Doing so provides four learned representations: real-life style, synthetic style, real-life content and synthetic content. Then, we combine them into feature representations from all combinations of style-content pairings across domains, and train a model on these combined representations to classify the content (i.e., labels) of a given datapoint in the style of another domain. We evaluate our method on real-life data using a variety of metrics to quantify the amount of information an attacker is able to recover. We show that our method prevents our model from overfitting to a small real-life training set, indicating that our method is an effective form of data augmentation, thereby making keystroke inference attacks more practical.
John Lim, Jan-Michael Frahm, Fabian Monrose
CODASPY2
2021 EgoGlass: Egocentric-View Human Pose Estimation From an Eyeglass Frame
abstract
We present a new approach, EgoGlass, towards egocentric motion-capture and human pose estimation. EgoGlass is a lightweight eyeglass frame with two cameras mounted on it. Our first contribution is a new egocentric motion-capture device that adds next to no extra burden on the user and a dataset of real people doing a diverse set of actions captured by EgoGlass. Second, we propose to utilize body part information for human pose detection - to help tackle the problems of limited body coverage and self-occlusions caused by the egocentric viewpoint and cameras’ proximity to the human body. We also propose a concept of pseudo-limb mask as an alternative for segmentation mask when ground truth segmentation mask is absent for egocentric images with real subject. We demonstrate that our method achieves better results than the counterpart method without body part information on our dataset. We also test our method on two existing egocentric datasets: xR-EgoPose and EgoCap. Our method achieves state-of-the-art results on xR-EgoPose and is on par with existing method for EgoCap without requiring temporal information or personalization for each individual user.
Dongxu Zhao 0001, Jisan Mahmud, Jan-Michael Frahm
3DV4
2021 RNNSLAM: Reconstructing the 3D colon to visualize missing regions during a colonoscopy
Ruibin Ma, Rui Wang 0071, Yubo Zhang 0004, Stephen M. Pizer, Sarah McGill, Julian G. Rosenman, Jan-Michael Frahm
Medical Image Anal.7
2020 Reducing Drift in Structure From Motion Using Extended Features
abstract
Low-frequency long-range errors (drift) are an endemic problem in 3D structure from motion, and can often hamper reasonable reconstructions of the scene. In this paper, we present a method to dramatically reduce scale and positional drift by using extended structural features such as planes and vanishing points. Unlike traditional feature matches, our extended features are able to span non-overlapping input images, and hence provide long-range constraints on the scale and shape of the reconstruction. We add these features as additional constraints to a state-of the-art global structure from motion algorithm and demonstrate that the added constraints enable the reconstruction of particularly drift-prone sequences such as long, low field-of-view videos without inertial measurements. Additionally, we provide an analysis of the drift-reducing capabilities of these constraints by evaluating on a synthetic dataset. Our structural features are able to significantly reduce drift for scenes that contain long-spanning man-made structures, such as aligned rows of windows or planar building facades.
Aleksander Holynski, David Geraghty, Jan-Michael Frahm, Chris Sweeney, Richard Szeliski
3DV3
2020 ViewSynth: Learning Local Features from Depth using View Synthesis
Jisan Mahmud, Rajat Vikram Singh, Peri Akiva, Spondon Kundu, Kuan-Chuan Peng, Jan-Michael Frahm
BMVC6
2020 Tangent Images for Mitigating Spherical Distortion
abstract
In this work, we propose "tangent images," a spherical image representation that facilitates transferable and scalable 360 degree computer vision. Inspired by techniques in cartography and computer graphics, we render a spherical image to a set of distortion-mitigated, locally-planar image grids tangent to a subdivided icosahedron. By varying the resolution of these grids independently of the subdivision level, we can effectively represent high resolution spherical images while still benefiting from the low-distortion icosahedral spherical approximation. We show that training standard convolutional neural networks on tangent images compares favorably to the many specialized spherical convolutional kernels that have been developed, while also scaling efficiently to handle significantly higher spherical resolutions. Furthermore, because our approach does not require specialized kernels, we show that we can transfer networks trained on perspective images to spherical data without fine-tuning and with limited performance drop-off. Finally, we demonstrate that tangent images can be used to improve the quality of sparse feature detection on spherical images, illustrating its usefulness for traditional computer vision tasks like structure-from-motion and SLAM.
Marc Eder, Mykhailo Shvets, John Lim, Jan-Michael Frahm
CVPR4
2020 Boundary-Aware 3D Building Reconstruction From a Single Overhead Image
abstract
We propose a boundary-aware multi-task deep-learning-based framework for fast 3D building modeling from a single overhead image. Unlike most existing techniques which rely on multiple images for 3D scene modeling, we seek to model the buildings in the scene from a single overhead image by jointly learning a modified signed distance function (SDF) from the building boundaries, a dense heightmap of the scene, and scene semantics. To jointly train for these tasks, we leverage pixel-wise semantic segmentation and normalized digital surface maps (nDSM) as supervision, in addition to labeled building outlines. At test time, buildings in the scene are automatically modeled in 3D using only an input overhead image. We demonstrate an increase in building modeling performance using a multi-feature network architecture that improves building outline detection by considering network features learned for the other jointly learned tasks. We also introduce a novel mechanism for robustly refining instance-specific building outlines using the learned modified SDF. We verify the effectiveness of our method on multiple large-scale satellite and aerial imagery datasets, where we obtain state-of-the-art performance in the 3D building reconstruction task.
Jisan Mahmud, True Price, Akash Bapat, Jan-Michael Frahm
CVPR4
2020 VPLNet: Deep Single View Normal Estimation With Vanishing Points and Lines
abstract
We present a novel single-view surface normal estimation method that combines traditional line and vanishing point analysis with a deep learning approach. Starting from a color image and a Manhattan line map, we use a deep neural network to regress on a dense normal map, and a dense Manhattan label map that identifies planar regions aligned with the Manhattan directions. We fuse the normal map and label map in a fully differentiable manner to produce a refined normal map as final output. To do so, we softly decompose the output into a Manhattan part and a non-Manhattan part. The Manhattan part is treated by discrete classification and vanishing points, while the non-Manhattan part is learned by direct supervision. Our method achieves state-of-the-art results on standard single-view normal estimation benchmarks. More importantly, we show that by using vanishing points and lines, our method has better generalization ability than existing works. In addition, we demonstrate how our surface normal network can improve the performance of depth estimation networks, both quantitatively and qualitatively, in particular, in 3D reconstructions of walls and other flat surfaces.
Rui Wang 0071, David Geraghty, Kevin Matzen, Richard Szeliski, Jan-Michael Frahm
CVPR5
2020 One shot 3D photography
abstract
3D photography is a new medium that allows viewers to more fully experience a captured moment. In this work, we refer to a 3D photo as one that displays parallax induced by moving the viewpoint (as opposed to a stereo pair with a fixed viewpoint). 3D photos are static in time, like traditional photos, but are displayed with interactive parallax on mobile or desktop screens, as well as on Virtual Reality devices, where viewing it also includes stereo. We present an end-to-end system for creating and viewing 3D photos, and the algorithmic and design choices therein. Our 3D photos are captured in a single shot and processed directly on a mobile device. The method starts by estimating depth from the 2D input image using a new monocular depth estimation network that is optimized for mobile devices. It performs competitively to the state-of-the-art, but has lower latency and peak memory consumption and uses an order of magnitude fewer parameters. The resulting depth is lifted to a layered depth image, and new geometry is synthesized in parallax regions. We synthesize color texture and structures in the parallax regions as well, using an inpainting network, also optimized for mobile devices, on the LDI directly. Finally, we convert the result into a mesh-based representation that can be efficiently transmitted and rendered even on low-end devices and over poor network connections. Altogether, the processing takes just a few seconds on a mobile device, and the result can be instantly viewed and shared. We perform extensive quantitative evaluation to validate our system and compare its new components against the current state-of-the-art.
Johannes Kopf 0001, Kevin Matzen, Suhib Alsisan, Ocean Quigley, Francis Ge, Yangming Chong, Josh Patterson, Jan-Michael Frahm, Matthew Yu, Peizhao Zhang, Peter Vajda, Ayush Saraf, Michael F. Cohen
ACM Trans. Graph.8
2019 The Domain Transform Solver
abstract
We present a novel framework for edge-aware optimization that is an order of magnitude faster than the state of the art while maintaining comparable results. Our key insight is that the optimization can be formulated by leveraging properties of the domain transform, a method for edge-aware filtering that defines a distance-preserving 1D mapping of the input space. This enables our method to improve performance for a wide variety of problems including stereo, depth super-resolution, render from defocus, colorization, and especially high-resolution depth filtering, while keeping the computational complexity linear in the number of pixels. Our method is highly parallelizable and adaptable, and it has demonstrable linear scalability with respect to image resolutions. We provide a comprehensive evaluation of our method w.r.t speed and accuracy for a variety of tasks.
Akash Bapat, Jan-Michael Frahm
CVPR2
2019 Recurrent Neural Network for (Un-)Supervised Learning of Monocular Video Visual Odometry and Depth
abstract
Deep learning-based, single-view depth estimation methods have recently shown highly promising results. However, such methods ignore one of the most important features for determining depth in the human vision system, which is motion. We propose a learning-based, multi-view dense depth map and odometry estimation method that uses Recurrent Neural Networks (RNN) and trains utilizing multi-view image reprojection and forward-backward flow-consistency losses. Our model can be trained in a supervised or even unsupervised mode. It is designed for depth and visual odometry estimation from video where the input frames are temporally correlated. However, it also generalizes to single-view depth estimation. Our method produces superior results to the state-of-the-art approaches for single-view and multi-view learning-based depth estimation on the KITTI driving dataset.
Rui Wang 0071, Stephen M. Pizer, Jan-Michael Frahm
CVPR3
2019 Real-Time 3D Reconstruction of Colonoscopic Surfaces for Determining Missing Regions
Ruibin Ma, Rui Wang 0071, Stephen M. Pizer, Julian G. Rosenman, Sarah McGill, Jan-Michael Frahm
MICCAI (5)6
2019 Re-Thinking CNN Frameworks for Time-Sensitive Autonomous-Driving Applications: Addressing an Industrial Challenge
abstract
Vision-based perception systems are crucial for profitable autonomous-driving vehicle products. High accuracy in such perception systems is being enabled by rapidly evolving convolution neural networks (CNNs). To achieve a better understanding of its surrounding environment, a vehicle must be provided with full coverage via multiple cameras. However, when processing multiple video streams, existing CNN frameworks often fail to provide enough inference performance, particularly on embedded hardware constrained by size, weight, and power limits. This paper presents the results of an industrial case study that was conducted to re-think the design of CNN software to better utilize available hardware resources. In this study, techniques such as parallelism, pipelining, and the merging of per-camera images into a single composite image were considered in the context of a Drive PX2 embedded hardware platform. The study identifies a combination of techniques that can be applied to increase throughput (number of simultaneous camera streams) without significantly increasing per-frame latency (camera to CNN output) or reducing per-stream accuracy.
Ming Yang 0036, Shige Wang, Joshua Bakita, Thanh Vu 0001, F. Donelson Smith, James H. Anderson, Jan-Michael Frahm
RTAS7
2018 Rolling Shutter and Radial Distortion Are Features for High Frame Rate Multi-Camera Tracking
abstract
Traditionally, camera-based tracking approaches have treated rolling shutter and radial distortion as imaging artifacts that have to be overcome and corrected for in order to apply standard camera models and scene reconstruction methods. In this paper, we introduce a novel multi-camera tracking approach that for the first time jointly leverages the information introduced by rolling shutter and radial distortion as a feature to achieve superior performance with respect to high-frequency camera pose estimation. In particular, our system is capable of attaining high tracking rates that were previously unachievable. Our approach explicitly leverages rolling shutter capture and radial distortion to process individual rows, rather than entire image frames, for accurate camera motion estimation. We estimate a per-row 6 DoF pose of a rolling shutter camera by tracking multiple points on a radially distorted row whose rays span a curved surface in 3D space. Although tracking systems for rolling shutter cameras exist, we are the first to leverage radial distortion to measure a per-row pose - enabling us to use less than half the number of cameras required by the previous state of the art. We validate our system on both synthetic and real imagery.
Akash Bapat, True Price, Jan-Michael Frahm
CVPR3
2018 Augmenting Crowd-Sourced 3D Reconstructions Using Semantic Detections
abstract
Image-based 3D reconstruction for Internet photo collections has become a robust technology to produce impressive virtual representations of real-world scenes. However, several fundamental challenges remain for Structure-from-Motion (SfM) pipelines, namely: the placement and reconstruction of transient objects only observed in single views, estimating the absolute scale of the scene, and (suprisingly often) recovering ground surfaces in the scene. We propose a method to jointly address these remaining open problems of SfM. In particular, we focus on detecting people in individual images and accurately placing them into an existing 3D model. As part of this placement, our method also estimates the absolute scale of the scene from object semantics, which in this case constitutes the height distribution of the population. Further, we obtain a smooth approximation of the ground surface and recover the gravity vector of the scene directly from the individual person detections. We demonstrate the results of our approach on a number of unordered Internet photo collections, and we quantitatively evaluate the obtained absolute scene scales.
True Price, Johannes L. Schönberger, Marc Pollefeys, Jan-Michael Frahm
CVPR5
2018 Hierarchy of Alternating Specialists for Scene Recognition
Hyo Jin Kim 0004, Jan-Michael Frahm
ECCV (11)2
2018 Improvement of Extrinsic Parameters from a Single Stereo Pair
abstract
In this paper, a novel algorithm for the automatic online improvement of the extrinsic camera parameters of a stereo image pair is introduced. To this end, the well-known dense stereo matching method PatchMatch stereo (PM) is extended for the pixelwise estimation of a discrepancy between the expected epipolar line and the actual correspondence. The availability of an initial guess of the camera parameters is assumed. Next, the estimated disparity map is filtered for highly stable and accurate correspondences that cover preferably the complete image. For this reason, we extend a quality estimation method adapted to Semi-Global Matching (SGM) derived disparity maps for general disparity maps. Finally, the set of stable and accurate correspondences from the disparity map is used for the estimation of the extrinsic camera parameters by means of the five-point algorithm in a RANSAC (random sample consensus) framework. Our algorithm can estimate optimized disparity maps and is able to adjust for errors in the relative camera pose. It can even correct epipolar errors of tens of pixels in highresolution images. We demonstrate that the proposed algorithm allows for robust and accurate estimation of the extrinsic camera parameters on datasets that provide weaklycalibrated image pairs.
Andreas Kuhn 0002, Lukas Roth, Jan-Michael Frahm, Helmut Mayer 0001
WACV3
2018 Dynamic Visual Sequence Prediction with Motion Flow Networks
abstract
We target the problem of synthesizing future motion sequences from a temporally ordered set of input images. Previous methods tackled this problem in two manners: predicting the future image pixel values and predicting the dense time-space trajectory of pixels. Towards this end, generative encoder-decoder networks have been widely adopted in both kinds of methods. However, pixel prediction with these networks has been shown to suffer from blurry outputs, since images are generated from scratch and there is no explicit enforcement of visual coherency. Alternately, crisp details can be achieved by transferring pixels from the input image through dense trajectory predictions, but this process requires pre-computed motion fields for training, which limit the learning ability for the neural networks. To synthesize realistic movement of objects under weak supervision (without pre-computed dense motion fields), we propose two novel network structures. Our first network encodes the input images as feature maps, and uses a decoder network to predict the future pixel correspondences for a series of subsequent time steps. The attained correspondence fields are then used to synthesize future views. Our second network focuses on human-centered capture by augmenting our framework to include sparse pose estimates [30] to guide our dense correspondence prediction. Compared with state-of-the-art pixel generating and dense trajectories predicting networks, our model performs better on synthetic as well as on real-world human body movement sequences.
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
WACV4
2018 Retweet Wars: Tweet Popularity Prediction via Dynamic Multimodal Regression
abstract
If a picture is worth a thousand words, then images should be utilized together with other available data modalities when predicting the virality of online posts, such as tweets. In this paper, we re-visit the tweet popularity prediction problem by considering all data modalities: tweet language semantics, embedded images, author' social relationships, and the diffusion process of tweets. To model the content of tweets, we propose a joint-embedding neural network that combines visual, textual, and social cues together. Such content features can be either used for prediction directly, or for pre-conditioning a 'dynamics RNN', which models the message propagation process. A novel Poisson regression loss is optimized to train the network. We demonstrate that content based features can be used to improve upon social features and dynamics features via our joint-embedding regression model. Our model outperforms the state-of-the-art on multiple large-scale real-world datasets collected from Twitter.
Ke Wang 0021, Mohit Bansal, Jan-Michael Frahm
WACV3
2018 Self-Expressive Dictionary Learning for Dynamic 3D Reconstruction
abstract
We target the problem of sparse 3D reconstruction of dynamic objects observed by multiple unsynchronized video cameras with unknown temporal overlap. To this end, we develop a framework to recover the unknown structure without sequencing information across video sequences. Our proposed compressed sensing framework poses the estimation of 3D structure as the problem of dictionary learning, where the dictionary is defined as an aggregation of the temporally varying 3D structures. Given the smooth motion of dynamic objects, we observe any element in the dictionary can be well approximated by a sparse linear combination of other elements in the same dictionary (i.e., self-expression). Our formulation optimizes a biconvex cost function that leverages a compressed sensing formulation and enforces both structural dependency coherence across video streams, as well as motion smoothness across estimates from common video sources. We further analyze the reconstructability of our approach under different capture scenarios, and its comparison and relation to existing methods. Experimental results on large amounts of synthetic data as well as real imagery demonstrate the effectiveness of our approach.
Enliang Zheng, Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
IEEE Trans. Pattern Anal. Mach. Intell.4
2018 Deep blending for free-viewpoint image-based rendering
abstract
Free-viewpoint image-based rendering (IBR) is a standing challenge. IBR methods combine warped versions of input photos to synthesize a novel view. The image quality of this combination is directly affected by geometric inaccuracies of multi-view stereo (MVS) reconstruction and by view- and image-dependent effects that produce artifacts when contributions from different input views are blended. We present a new deep learning approach to blending for IBR, in which we use held-out real image data to learn blending weights to combine input photo contributions. Our Deep Blending method requires us to address several challenges to achieve our goal of interactive free-viewpoint IBR navigation. We first need to provide sufficiently accurate geometry so the Convolutional Neural Network (CNN) can succeed in finding correct blending weights. We do this by combining two different MVS reconstructions with complementary accuracy vs. completeness tradeoffs. To tightly integrate learning in an interactive IBR system, we need to adapt our rendering algorithm to produce a fixed number of input layers that can then be blended by the CNN. We generate training data with a variety of captured scenes, using each input photo as ground truth in a held-out approach. We also design the network architecture and the training loss to provide high quality novel view synthesis, while reducing temporal flickering artifacts. Our results demonstrate free-viewpoint IBR in a wide variety of scenes, clearly surpassing previous methods in visual quality, especially when moving far from the input cameras.
Peter Hedman, Julien Philip, True Price, Jan-Michael Frahm, George Drettakis, Gabriel J. Brostow
ACM Trans. Graph.4
2018 Towards Fully Mobile 3D Face, Body, and Environment Capture Using Only Head-worn Cameras
abstract
We propose a new approach for 3D reconstruction of dynamic indoor and outdoor scenes in everyday environments, leveraging only cameras worn by a user. This approach allows 3D reconstruction of experiences at any location and virtual tours from anywhere. The key innovation of the proposed ego-centric reconstruction system is to capture the wearer's body pose and facial expression from near-body views, e.g. cameras on the user's glasses, and to capture the surrounding environment using outward-facing views. The main challenge of the ego-centric reconstruction, however, is the poor coverage of the near-body views - that is, the user's body and face are observed from vantage points that are convenient for wear but inconvenient for capture. To overcome these challenges, we propose a parametric-model-based approach to user motion estimation. This approach utilizes convolutional neural networks (CNNs) for near-view body pose estimation, and we introduce a CNN-based approach for facial expression estimation that combines audio and video. For each time-point during capture, the intermediate model-based reconstructions from these systems are used to re-target a high-fidelity pre-scanned model of the user. We demonstrate that the proposed self-sufficient, head-worn capture system is capable of reconstructing the wearer's movements and their surrounding environment in both indoor and outdoor situations without any additional views. As a proof of concept, we show how the resulting 3D-plus-time reconstruction can be immersively experienced within a virtual reality system (e.g., the HTC Vive). We expect that the size of the proposed egocentric capture-and-reconstruction system will eventually be reduced to fit within future AR glasses, and will be widely useful for immersive 3D telepresence, virtual tours, and general use-anywhere 3D content creation.
Young-Woon Cha, True Price, Xinran Lu, Nicholas Rewkowski, Rohan Chabra, Zihe Qin, Hyounghun Kim, Zhaoqi Su, Yebin Liu, Adrian Ilie, Andrei State, Zhenlin Xu, Jan-Michael Frahm, Henry Fuchs
IEEE Trans. Vis. Comput. Graph.14
2017 Fast and Accurate Satellite Multi-view Stereo Using Edge-Aware Interpolation
abstract
In this paper, we propose a fast and accurate approach for 3D reconstructions from satellite images. Compared with traditional images, satellite imagery features enormous pixel count, inaccurate camera calibration, and low ground sampling rate, all of which makes multi-view stereo for satellite images more challenging. Our approach first computes sparse but reliable 2D feature matches between image pairs. Such feature matches are used to compensate the extrinsic calibration errors. Preliminary dense correspondences are obtained via edge-aware interpolation of sparse feature matches. We rely on fast bilateral smoothing to refine such initial dense matches, which greatly improves the computational efficiency of our method. The smoothed dense correspondences and the refined camera model are then used to obtain dense 3D point clouds via triangulation. Our proposed method outperforms state-of-the-art baseline methods in both efficiency and accuracy on real-world datasets.
Ke Wang 0021, Jan-Michael Frahm
3DV2
2017 Single View Parametric Building Reconstruction from Satellite Imagery
abstract
Satellite images have broad coverage, thus are ideal for large-scale urban reconstruction tasks. However, their low ground sampling resolution posed great challenges in using traditional volumetric or stereo methods to perform 3D reconstructions. In this paper, we propose a novel deep learning based approach to perform single-view parametric reconstructions from satellite imagery. By parametrizing buildings as 3D cuboids, our method extends object detection systems to simultaneously localize buildings and directly fit parametric models for each identified building. We utilize geo-registered GIS vector maps and Lidar data as supervision to train the network. Especially, we deconvolve the feature maps and combine convolutional feature maps at different stages of the network to deal with the heavily cluttered but small in size building instances from satellite imagery. We further enforce physical constraints that building cannot overlap by predicting building boundaries using a separate fully convolutional network. We demonstrate the effectiveness of our proposed methods on real-world data.
Ke Wang 0021, Jan-Michael Frahm
3DV2
2017 Learned Contextual Feature Reweighting for Image Geo-Localization
abstract
We address the problem of large scale image geo-localization where the location of an image is estimated by identifying geo-tagged reference images depicting the same place. We propose a novel model for learning image representations that integrates context-aware feature reweighting in order to effectively focus on regions that positively contribute to geo-localization. In particular, we introduce a Contextual Reweighting Network (CRN) that predicts the importance of each region in the feature map based on the image context. Our model is learned end-to-end for the image geo-localization task, and requires no annotation other than image geo-tags for training. In experimental results, the proposed approach significantly outperforms the previous state-of-the-art on the standard geo-localization benchmark datasets. We also demonstrate that our CRN discovers task-relevant contexts without any additional supervision.
Hyo Jin Kim 0004, Enrique Dunn, Jan-Michael Frahm
CVPR3
2017 The Misty Three Point Algorithm for Relative Pose
abstract
There is a significant interest in scene reconstruction from underwater images given its utility for oceanic research and for recreational image manipulation. In this paper we propose a novel algorithm for two view camera motion estimation for underwater imagery. Our method leverages the constraints provided by the attenuation properties of water and its effects on the appearance of the color to determine the depth difference of a point with respect to the two observing views of the underwater cameras. Additionally, we propose an algorithm, leveraging the depth differences of three such observed points, to estimate the relative pose of the cameras. Given the unknown underwater attenuation coefficients, our method estimates the relative motion up to scale. The results are represented as a generalized camera. We evaluate our method on both real data and simulated data.
Tobias Palmér, Kalle Åström, Jan-Michael Frahm
CVPR3
2016 A Vote-and-Verify Strategy for Fast Spatial Verification in Image Retrieval
Johannes L. Schönberger, True Price, Torsten Sattler, Jan-Michael Frahm, Marc Pollefeys
ACCV (1)4
2016 Bringing 3D Models Together: Mining Video Liaisons in Crowdsourced Reconstructions
Ke Wang 0021, Enrique Dunn, Mikel Rodriguez, Jan-Michael Frahm
ACCV (4)4
2016 From Dusk Till Dawn: Modeling in the Dark
abstract
Internet photo collections naturally contain a large variety of illumination conditions, with the largest difference between day and night images. Current modeling techniques do not embrace the broad illumination range often leading to reconstruction failure or severe artifacts. We present an algorithm that leverages the appearance variety to obtain more complete and accurate scene geometry along with consistent multi-illumination appearance information. The proposed method relies on automatic scene appearance grouping, which is used to obtain separate dense 3D models. Subsequent model fusion combines the separate models into a complete and accurate reconstruction of the scene. In addition, we propose a method to derive the appearance information for the model under the different illumination conditions, even for scene parts that are not observed under one illumination condition. To achieve this, we develop a cross-illumination color transfer technique. We evaluate our method on a large variety of landmarks from across Europe reconstructed from a database of 7.4M images.
Filip Radenovic, Johannes L. Schönberger, Dinghuang Ji, Jan-Michael Frahm, Ondrej Chum, Jiri Matas
CVPR4
2016 Structure-from-Motion Revisited
abstract
Incremental Structure-from-Motion is a prevalent strategy for 3D reconstruction from unordered image collections. While incremental reconstruction systems have tremendously advanced in all regards, robustness, accuracy, completeness, and scalability remain the key problems towards building a truly general-purpose pipeline. We propose a new SfM technique that improves upon the state of the art to make a further step towards this ultimate goal. The full reconstruction pipeline is released to the public as an open-source implementation.
Johannes L. Schönberger, Jan-Michael Frahm
CVPR2
2016 Years-Long Binary Image Broadcast Using Bluetooth Low Energy Beacons
abstract
This paper describes the first 'image beacon' system that is capable of broadcasting binary images over a very long period (years, as opposed to days or weeks) using a set of cheap, low-power, memory-constrained Bluetooth Low Energy (BLE) beacon devices. We design a patch-based image encoding algorithm to produce encoded images of reasonably high quality, having sizes of as low as 16 bytes -- without any prior knowledge of the test images. We test our system with different types of images that contain hand-written alphanumeric characters, geometric shapes, and arbitrary binary images having complex shapes and curves. We empirically determine the tradeoffs between the system lifetime and the quality of broadcasted images, and determine an optimal set of parameters for our system, under user-specified constraints such as the number of available beacon devices, maximum latency, and life expectancy. We develop a smartphone application that takes an image and user-requirements as inputs, shows previews of different quality output images, writes the encoded image into a set of beacons, and reads the broadcasted image back. Our evaluation shows that a set of 2 -- 3 beacons is capable of broadcasting high-quality images (75% -- 90% structurally similar to original images) for a year-long continuous broadcasting, and both the lifetime and the image quality improve when more beacons are used.
Chong Shao, Shahriar Nirjon, Jan-Michael Frahm
DCOSS3
2016 Indoor-Outdoor 3D Reconstruction Alignment
Andrea Cohen, Johannes L. Schönberger, Pablo Speciale, Torsten Sattler, Jan-Michael Frahm, Marc Pollefeys
ECCV (3)5
2016 Spatio-Temporally Consistent Correspondence for Dense Dynamic Scene Modeling
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
ECCV (6)3
2016 Pixelwise View Selection for Unstructured Multi-View Stereo
Johannes L. Schönberger, Enliang Zheng, Jan-Michael Frahm, Marc Pollefeys
ECCV (3)3
2016 Virtual U: Defeating Face Liveness Detection by Building Virtual Models from Your Public Photos
Yi Xu 0006, True Price, Jan-Michael Frahm, Fabian Monrose
USENIX Security Symposium3
2016 Efficient joint stereo estimation and land usage classification for multiview satellite data
abstract
We propose an efficient algorithm to jointly estimate geometry and semantics for a given geographical region observed by multiple satellite images. Our joint estimation leverages an efficient PatchMatch inference framework defined over lattice discretization of the environment. Our cost function relies on the local planarity assumption to model scene geometry and neural network classification to determine semantic (e.g. land use) labels for geometric structures. By utilizing the commonly available direct (i.e. space to image) rational polynomial coefficients (RPC) satellite camera models, our approach effectively circumvents the need for estimating or refining inverse RPC models. Experiments illustrate both the computational efficiency and high quality scene geometry estimates attained by our approach for satellite imagery. To further illustrate the generality of our representation and inference framework, experiments on standard benchmarks for ground-level imagery are also included.
Ke Wang 0021, Craig Stutts, Enrique Dunn, Jan-Michael Frahm
WACV4
2016 Towards Kilo-Hertz 6-DoF Visual Tracking Using an Egocentric Cluster of Rolling Shutter Cameras
abstract
To maintain a reliable registration of the virtual world with the real world, augmented reality (AR) applications require highly accurate, low-latency tracking of the device. In this paper, we propose a novel method for performing this fast 6-DOF head pose tracking using a cluster of rolling shutter cameras. The key idea is that a rolling shutter camera works by capturing the rows of an image in rapid succession, essentially acting as a high-frequency 1D image sensor. By integrating multiple rolling shutter cameras on the AR device, our tracker is able to perform 6-DOF markerless tracking in a static indoor environment with minimal latency. Compared to state-of-the-art tracking systems, this tracking approach performs at significantly higher frequency, and it works in generalized environments. To demonstrate the feasibility of our system, we present thorough evaluations on synthetically generated data with tracking frequencies reaching 56.7 kHz. We further validate the method's accuracy on real-world images collected from a prototype of our tracking system against ground truth data using standard commodity GoPro cameras capturing at 120 Hz frame rate.
Akash Bapat, Enrique Dunn, Jan-Michael Frahm
IEEE Trans. Vis. Comput. Graph.3
2015 Reconstructing the world* in six days
abstract
We propose a novel, large-scale, structure-from-motion framework that advances the state of the art in data scalability from city-scale modeling (millions of images) to world-scale modeling (several tens of millions of images) using just a single computer. The main enabling technology is the use of a streaming-based framework for connected component discovery. Moreover, our system employs an adaptive, online, iconic image clustering approach based on an augmented bag-of-words representation, in order to balance the goals of registration, comprehensiveness, and data compactness. We demonstrate our proposal by operating on a recent publicly available 100 million image crowd-sourced photo collection containing images geographically distributed throughout the entire world. Results illustrate that our streaming-based approach does not compromise model completeness, but achieves unprecedented levels of efficiency and scalability.
Jared Heinly, Johannes L. Schönberger, Enrique Dunn, Jan-Michael Frahm
CVPR4
2015 Adaptive eye-camera calibration for head-worn devices
abstract
We present a novel, continuous, locally optimal calibration scheme for use with head-worn devices. Current calibration schemes solve for a globally optimal model of the eye-device transformation by performing calibration on a per-user or once-per-use basis. However, these calibration schemes are impractical for real-world applications because they do not account for changes in calibration during the time of use. Our calibration scheme allows a head-worn device to calculate a locally optimal eye-device transformation on demand by computing an optimal model from a local window of previous frames. By leveraging naturally occurring interest regions within the user's environment, our system can calibrate itself without the user's active participation. Experimental results demonstrate that our proposed calibration scheme outperforms the existing state of the art systems while being significantly less restrictive to the user and the environment.
David Perra, Rohit Kumar Gupta, Jan-Michael Frahm
CVPR3
2015 PAIGE: PAirwise Image Geometry Encoding for improved efficiency in Structure-from-Motion
abstract
Large-scale Structure-from-Motion systems typically spend major computational effort on pairwise image matching and geometric verification in order to discover connected components in large-scale, unordered image collections. In recent years, the research community has spent significant effort on improving the efficiency of this stage. In this paper, we present a comprehensive overview of various state-of-the-art methods, evaluating and analyzing their performance. Based on the insights of this evaluation, we propose a learning-based approach, the PAirwise Image Geometry Encoding (PAIGE), to efficiently identify image pairs with scene overlap without the need to perform exhaustive putative matching and geometric verification. PAIGE achieves state-of-the-art performance and integrates well into existing Structure-from-Motion pipelines.
Johannes L. Schönberger, Alexander C. Berg, Jan-Michael Frahm
CVPR3
2015 From single image query to detailed 3D reconstruction
abstract
Structure-from-Motion for unordered image collections has significantly advanced in scale over the last decade. This impressive progress can be in part attributed to the introduction of efficient retrieval methods for those systems. While this boosts scalability, it also limits the amount of detail that the large-scale reconstruction systems are able to produce. In this paper, we propose a joint reconstruction and retrieval system that maintains the scalability of large-scale Structure-from-Motion systems while also recovering the often lost ability of reconstructing fine details of the scene. We demonstrate our proposed method on a large-scale dataset of 7.4 million images downloaded from the Internet.
Johannes L. Schönberger, Filip Radenovic, Ondrej Chum, Jan-Michael Frahm
CVPR4
2015 Synthesizing Illumination Mosaics from Internet Photo-Collections
abstract
We propose a framework for the automatic creation of time-lapse mosaics of a given scene. We achieve this by leveraging the illumination variations captured in Internet photo-collections. In order to depict and characterize the illumination spectrum of a scene, our method relies on building discrete representations of the image appearance space through connectivity graphs defined over a pairwise image distance function. The smooth appearance transitions are found as the shortest path in the similarity graph among images, and robust image alignment is achieved by leveraging scene semantics, multi-view geometry, and image warping techniques. The attained results present an insightful and compact visualization of the scene illuminations captured in crowd-sourced imagery.
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
ICCV3
2015 Predicting Good Features for Image Geo-Localization Using Per-Bundle VLAD
abstract
We address the problem of recognizing a place depicted in a query image by using a large database of geo-tagged images at a city-scale. In particular, we discover features that are useful for recognizing a place in a data-driven manner, and use this knowledge to predict useful features in a query image prior to the geo-localization process. This allows us to achieve better performance while reducing the number of features. Also, for both learning to predict features and retrieving geo-tagged images from the database, we propose per-bundle vector of locally aggregated descriptors (PBVLAD), where each maximally stable region is described by a vector of locally aggregated descriptors (VLAD) on multiple scale-invariant features detected within the region. Experimental results show the proposed approach achieves a significant improvement over other baseline methods.
Hyo Jin Kim 0004, Enrique Dunn, Jan-Michael Frahm
ICCV3
2015 Sparse Dynamic 3D Reconstruction from Unsynchronized Videos
abstract
We target the sparse 3D reconstruction of dynamic objects observed by multiple unsynchronized video cameras with unknown temporal overlap. To this end, we develop a framework to recover the unknown structure without sequencing information across video sequences. Our proposed compressed sensing framework poses the estimation of 3D structure as the problem of dictionary learning. Moreover, we define our dictionary as the temporally varying 3D structure, while we define local sequencing information in terms of the sparse coefficients describing a locally linear 3D structural interpolation. Our formulation optimizes a biconvex cost function that leverages a compressed sensing formulation and enforces both structural dependency coherence across video streams, as well as motion smoothness across estimates from common video sources. Experimental results demonstrate the effectiveness of our approach in both synthetic data and captured imagery.
Enliang Zheng, Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
ICCV4
2015 Minimal Solvers for 3D Geometry from Satellite Imagery
abstract
We propose two novel minimal solvers which advance the state of the art in satellite imagery processing. Our methods are efficient and do not rely on the prior existence of complex inverse mapping functions to correlate 2D image coordinates and 3D terrain. Our first solver improves on the stereo correspondence problem for satellite imagery, in that we provide an exact image-to-object space mapping (where prior methods were inaccurate). Our second solver provides a novel mechanism for 3D point triangulation, which has improved robustness and accuracy over prior techniques. Given the usefulness and ubiquity of satellite imagery, our proposed methods allow for improved results in a variety of existing and future applications.
Enliang Zheng, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm
ICCV4
2014 Recovering Correct Reconstructions from Indistinguishable Geometry
abstract
Structure-from-motion (SFM) is widely utilized to generate 3D reconstructions from unordered photo-collections. However, in the presence of non unique, symmetric, or otherwise indistinguishable structure, SFM techniques often incorrectly reconstruct the final model. We propose a method that not only determines if an error is present, but automatically corrects the error in order to produce a correct representation of the scene. We find that by exploiting the co-occurrence information present in the scene's geometry, we can successfully isolate the 3D points causing the incorrect result. This allows us to split an incorrect reconstruction into error-free sub-models that we then correctly merge back together. Our experimental results show that our technique is efficient, robust to a variety of scenes, and outperforms existing methods.
Jared Heinly, Enrique Dunn, Jan-Michael Frahm
3DV3
2014 Cloud-scale Image Compression Through Content Deduplication
David Perra, Jan-Michael Frahm
BMVC2
2014 Watching the Watchers: Automatically Inferring TV Content From Outdoor Light Effusions
abstract
The flickering lights of content playing on TV screens in our living rooms are an all too familiar sight at night --- and one that many of us have paid little attention to with regards to the amount of information these diffusions may leak to an inquisitive outsider. In this paper, we introduce an attack that exploits the emanations of changes in light (e.g., as seen through the windows and recorded over 70 meters away) to reveal the programs we watch. Our empirical results show that the attack is surprisingly robust to a variety of noise signals that occur in real-world situations, and moreover, can successfully identify the content being watched among a reference library of tens of thousands of videos within several seconds. The robustness and efficiency of the attack can be attributed to the use of novel feature sets and an elegant online algorithm for performing index-based matches.
Yi Xu 0006, Jan-Michael Frahm, Fabian Monrose
CCS2
2014 Stereo under Sequential Optimal Sampling: A Statistical Analysis Framework for Search Space Reduction
abstract
We develop a sequential optimal sampling framework for stereo disparity estimation by adapting the Sequential Probability Ratio Test (SPRT) model. We operate over local image neighborhoods by iteratively estimating single pixel disparity values until sufficient evidence has been gathered to either validate or contradict the current hypothesis regarding local scene structure. The output of our sampling is a set of sampled pixel positions along with a robust and compact estimate of the set of disparities contained within a given region. We further propose an efficient plane propagation mechanism that leverages the pre-computed sampling positions and the local structure model described by the reduced local disparity set. Our sampling framework is a general pre-processing mechanism aimed at reducing computational complexity of disparity search algorithms by ascertaining a reduced set of disparity hypotheses for each pixel. Experiments demonstrate the effectiveness of the proposed approach when compared to state of the art methods.
Yilin Wang 0001, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm
CVPR4
2014 PatchMatch Based Joint View Selection and Depthmap Estimation
abstract
We propose a multi-view depthmap estimation approach aimed at adaptively ascertaining the pixel level data associations between a reference image and all the elements of a source image set. Namely, we address the question, what aggregation subset of the source image set should we use to estimate the depth of a particular pixel in the reference image? We pose the problem within a probabilistic framework that jointly models pixel-level view selection and depthmap estimation given the local pairwise image photoconsistency. The corresponding graphical model is solved by EM-based view selection probability inference and PatchMatch-like depth sampling and propagation. Experimental results on standard multi-view benchmarks convey the state-of-the art estimation accuracy afforded by mitigating spurious pixel level data associations. Additionally, experiments on large Internet crowd sourced data demonstrate the robustness of our approach against unstructured and heterogeneous image capture characteristics. Moreover, the linear computational and storage requirements of our formulation, as well as its inherent parallelism, enables an efficient and scalable GPU-based implementation.
Enliang Zheng, Enrique Dunn, Vladimir Jojic, Jan-Michael Frahm
CVPR4
2014 Correcting for Duplicate Scene Structure in Sparse 3D Reconstruction
Jared Heinly, Enrique Dunn, Jan-Michael Frahm
ECCV (4)3
2014 3D Reconstruction of Dynamic Textures in Crowd Sourced Data
Dinghuang Ji, Enrique Dunn, Jan-Michael Frahm
ECCV (1)3
2014 Joint Object Class Sequencing and Trajectory Triangulation (JOST)
Enliang Zheng, Ke Wang 0021, Enrique Dunn, Jan-Michael Frahm
ECCV (7)4
2014 P-HRTF: Efficient personalized HRTF computation for high-fidelity spatial sound
abstract
Accurate rendering of 3D spatial audio for interactive virtual auditory displays requires the use of personalized head-related transfer functions (HRTFs). We present a new approach to compute personalized HRTFs for any individual using a method that combines state-of-the-art image-based 3D modeling with an efficient numerical simulation pipeline. Our 3D modeling framework enables capture of the listener's head and torso using consumer-grade digital cameras to estimate a high-resolution non-parametric surface representation of the head, including the extended vicinity of the listener's ear. We leverage sparse structure from motion and dense surface reconstruction techniques to generate a 3D mesh. This mesh is used as input to a numeric sound propagation solver, which uses acoustic reciprocity and Kirchhoff surface integral representation to efficiently compute an individual's personalized HRTF. The overall computation takes tens of minutes on multi-core desktop machine. We have used our approach to compute the personalized HRTFs of few individuals, and we present our preliminary evaluation here. To the best of our knowledge, this is the first commodity technique that can be used to compute personalized HRTFs in a lab or home setting.
Alok Meshram, Ravish Mehra, Hongsheng Yang, Enrique Dunn, Jan-Michael Frahm, Dinesh Manocha
ISMAR5
2014 Rotation estimation from cloud tracking
abstract
We address the problem of online relative orientation estimation from streaming video captured by a sky-facing camera on a mobile device. Namely, we rely on the detection and tracking of visual features attained from cloud structures. Our proposed method achieves robust and efficient operation by combining realtime visual odometry modules, learning based feature classification, and Kalman filtering within a robustness-driven data management framework, while achieving framerate processing on a mobile device. The relatively large 3D distance between the camera and the observed cloud features is leveraged to simplify our processing pipeline. First, as an efficiency driven optimization, we adopt a homography based motion model and focus on estimating relative rotations across adjacent keyframes. To this end, we rely on efficient feature extraction, KLT tracking, and RANSAC based model fitting. Second, to ensure the validity of our simplified motion model, we segregate detected cloud features from scene features through SVM classification. Finally, to make tracking more robust, we employ predictive Kalman filtering to enable feature persistence through temporary occlusions and manage feature spatial distribution to foster tracking robustness. Results exemplify the accuracy and robustness of the proposed approach and highlight its potential as a passive orientation sensor.
Sangwoo Cho, Enrique Dunn, Jan-Michael Frahm
WACV3
2014 Combining semantic scene priors and haze removal for single image depth estimation
abstract
We consider the problem of estimating the relative depth of a scene from a monocular image. The dark channel prior, used as a statistical observation of haze free images, has been previously leveraged for haze removal and relative depth estimation tasks. However, as a local measure, it fails to account for higher order semantic relationship among scene elements. We propose a dual channel prior used for identifying pixels that are unlikely to comply with the dark channel assumption, leading to erroneous depth estimates. We further leverage semantic segmentation information and patch match label propagation to enforce semantically consistent geometric priors. Experiments illustrate the quantitative and qualitative advantages of our approach when compared to state of the art methods.
Ke Wang 0021, Enrique Dunn, Joseph Tighe, Jan-Michael Frahm
WACV4
2014 Security Analysis and Related Usability of Motion-Based CAPTCHAs: Decoding Codewords in Motion
abstract
We explore the robustness and usability of moving-image object recognition (video) CAPTCHAS, designing and implementing automated attacks based on computer vision techniques. Our approach is suitable for broad classes of moving-image CAPTCHAS involving rigid objects. We first present an attack that defeats instances of such a CAPTCHA (NuCaptcha) representing the state-of-the-art, involving dynamic text strings called codewords. We then consider design modifications to mitigate the attacks (e.g., overlapping characters more closely, randomly changing the font of individual characters, or even randomly varying the number of characters in the codeword). We implement the modified CAPTCHAS and test if designs modified for greater robustness maintain usability. Our lab-based studies show that the modified captchas fail to offer viable usability, even when the captcha strength is reduced below acceptable targets. Worse yet, our GPU-based implementation shows that our automated approach can decode these captchas faster than humans can, and we can do so at a relatively low cost of roughly 50 cents per 1,000 captchas solved based on Amazon EC2 rates circa 2012. To further demonstrate the challenges in designing usable captchas, we also implement and test another variant of moving text strings using the known emerging images concept. This variant is resilient to our attacks and also offers similar usability to commercially available approaches. We explain why fundamental elements of the emerging images idea resist our current attack where others fail.
Yi Xu 0006, Gerardo Reynaga, Sonia Chiasson, Jan-Michael Frahm, Fabian Monrose, Paul C. van Oorschot
IEEE Trans. Dependable Secur. Comput.4
2014 Personal Photograph Enhancement Using Internet Photo Collections
abstract
Given the growth of Internet photo collections, we now have a visual index of all major cities and tourist sites in the world. However, it is still a difficult task to capture that perfect shot with your own camera when visiting these places, especially when your camera itself has limitations, such as a limited field of view. In this paper, we propose a framework to overcome the imperfections of personal photographs of tourist sites using the rich information provided by large-scale Internet photo collections. Our method deploys state-of-the-art techniques for constructing initial 3D models from photo collections. The same techniques are then used to register personal photographs to these models, allowing us to augment personal 2D images with 3D information. This strong available scene prior allows us to address a number of traditionally challenging image enhancement techniques and achieve high-quality results using simple and robust algorithms. Specifically, we demonstrate automatic foreground segmentation, mono-to-stereo conversion, field-of-view expansion, photometric enhancement, and additionally automatic annotation with geolocation and tags. Our method clearly demonstrates some possible benefits of employing the rich information contained in online photo databases to efficiently enhance and augment one's own personal photographs.
Jizhou Gao, Oliver Wang, Pierre Fite Georgel, Ruigang Yang, James Davis 0001, Jan-Michael Frahm, Marc Pollefeys
IEEE Trans. Vis. Comput. Graph.7
2013 Seeing double: reconstructing obscured typed input from repeated compromising reflections
abstract
Of late, threats enabled by the ubiquitous use of mobile devices have drawn much interest from the research community. However, prior threats all suffer from a similar, and profound, weakness - namely the requirement that the adversary is either within visual range of the victim (e.g., to ensure that the pop-out events in reflections in the victim's sunglasses can be discerned) or is close enough to the target to avoid the use of expensive telescopes. In this paper, we broaden the scope of the attacks by relaxing these requirements and show that breaches of privacy are possible even when the adversary is around a corner. The approach we take overcomes challenges posed by low image resolution by extending computer vision methods to operate on small, high-noise, images. Moreover, our work is applicable to all types of keyboards because of a novel application of fingertip motion analysis for key-press detection. In doing so, we are also able to exploit reflections in the eyeball of the user or even repeated reflections (i.e., a reflection of a reflection of the mobile device in the eyeball of the user). Our empirical results show that we can perform these attacks with high accuracy, and can do so in scenarios that aptly demonstrate the realism of this threat.
Yi Xu 0006, Jared Heinly, Andrew M. White 0002, Fabian Monrose, Jan-Michael Frahm
CCS5
2013 Scanning and tracking dynamic objects with commodity depth cameras
abstract
The 3D data collected using state-of-the-art algorithms often suffers from various problems, such as incompletion and inaccuracy. Using temporal information has been proven effective for improving the reconstruction quality; for example, KinectFusion [21] shows significant improvements for static scenes. In this work, we present a system that uses commodity depth and color cameras, such as Microsoft Kinects, to fuse the 3D data captured over time for dynamic objects to build a complete and accurate model, and then tracks the model to match later observations. The key ingredients of our system include a nonrigid matching algorithm that aligns 3D observations of dynamic objects by using both geometry and texture measurements, and a volumetric fusion algorithm that fuses noisy 3D data. We demonstrate that the quality of the model improves dramatically by fusing a sequence of noisy and incomplete depth data of human and that by deforming this fused model to later observations, noise-and-hole-free 3D models are generated for the human moving freely.
Mingsong Dou, Henry Fuchs, Jan-Michael Frahm
ISMAR3
2013 USAC: A Universal Framework for Random Sample Consensus
abstract
A computational problem that arises frequently in computer vision is that of estimating the parameters of a model from data that have been contaminated by noise and outliers. More generally, any practical system that seeks to estimate quantities from noisy data measurements must have at its core some means of dealing with data contamination. The random sample consensus (RANSAC) algorithm is one of the most popular tools for robust estimation. Recent years have seen an explosion of activity in this area, leading to the development of a number of techniques that improve upon the efficiency and robustness of the basic RANSAC algorithm. In this paper, we present a comprehensive overview of recent research in RANSAC-based robust estimation by analyzing and comparing various approaches that have been explored over the years. We provide a common context for this analysis by introducing a new framework for robust estimation, which we call Universal RANSAC (USAC). USAC extends the simple hypothesize-and-verify structure of standard RANSAC to incorporate a number of important practical and computational considerations. In addition, we provide a general-purpose C++ software library that implements the USAC framework by leveraging state-of-the-art algorithms for the various modules. This implementation thus addresses many of the limitations of standard RANSAC within a single unified package. We benchmark the performance of the algorithm on a large collection of estimation problems. The implementation we provide can be used by researchers either as a stand-alone tool for robust estimation or as a benchmark for evaluating new techniques.
Rahul Raguram, Ondrej Chum, Marc Pollefeys, Jiri Matas, Jan-Michael Frahm
IEEE Trans. Pattern Anal. Mach. Intell.5
2013 On the Privacy Risks of Virtual Keyboards: Automatic Reconstruction of Typed Input from Compromising Reflections
abstract
We investigate the implications of the ubiquity of personal mobile devices and reveal new techniques for compromising the privacy of users typing on virtual keyboards. Specifically, we show that so-called compromising reflections (in, for example, a victim's sunglasses) of a device's screen are sufficient to enable automated reconstruction, from video, of text typed on a virtual keyboard. Through the use of advanced computer vision and machine learning techniques, we are able to operate under extremely realistic threat models, in real-world operating conditions, which are far beyond the range of more traditional OCR-based attacks. In particular, our system does not require expensive and bulky telescopic lenses: rather, we make use of off-the-shelf, handheld video cameras. In addition, we make no limiting assumptions about the motion of the phone or of the camera, nor the typing style of the user, and are able to reconstruct accurate transcripts of recorded input, even when using footage captured in challenging environments (e.g., on a moving bus). To further underscore the extent of this threat, our system is able to achieve accurate results even at very large distances-up to 61 m for direct surveillance, and 12 m for sunglass reflections. We believe these results highlight the importance of adjusting privacy expectations in response to emerging technologies.
Rahul Raguram, Andrew M. White 0002, Yi Xu 0006, Jan-Michael Frahm, Pierre Fite Georgel, Fabian Monrose
IEEE Trans. Dependable Secur. Comput.4
2012 Improved Geometric Verification for Large Scale Landmark Image Collections
abstract
In this work, we address the issue of geometric verification, with a focus on modeling large-scale landmark image collections gathered from the internet. In particular, we show that we can compute and learn descriptive statistics pertaining to the image collection by leveraging information that arises as a by-product of the matching and verification stages. Our approach is based on the intuition that validating numerous image pairs of the same geometric scene structures quickly reveals useful information about two aspects of the image collection: (a) the reliability of individual visual words and (b) the appearance of landmarks in the image collection. Both of these sources of information can then be used to drive any subsequent processing, thus allowing the system to bootstrap itself. While current techniques make use of dedicated training/preprocessing stages, our approach elegantly integrates into the standard geometric verification pipeline, by simply leveraging the information revealed during the verification stage. The main result of this work is that this unsupervised “learning-as-you-go ” approach significantly improves performance; our experiments demonstrate significant improvements in efficiency and completeness over standard techniques.
Rahul Raguram, Joseph Tighe, Jan-Michael Frahm
BMVC3
2012 Efficient and Scalable Depthmap Fusion
Enliang Zheng, Enrique Dunn, Rahul Raguram, Jan-Michael Frahm
BMVC4
2012 Comparative Evaluation of Binary Features
Jared Heinly, Enrique Dunn, Jan-Michael Frahm
ECCV (2)3
2012 Security and Usability Challenges of Moving-Object CAPTCHAs: Decoding Codewords in Motion
Yi Xu 0006, Gerardo Reynaga, Sonia Chiasson, Jan-Michael Frahm, Fabian Monrose, Paul C. van Oorschot
USENIX Security Symposium4
2012 Room-sized informal telepresence system
abstract
We present a room-sized telepresence system for informal gatherings rather than conventional meetings. Unlike conventional systems which constrain participants to sit in fixed positions, our system aims to facilitate casual conversations between people in two sites. The system consists of a wall of large flat displays at each of the two sites, showing a panorama of the remote scene, constructed from a multiplicity of color and depth cameras. The main contribution of this paper is a solution that ameliorates the eye contact problem during conversation in typical scenarios while still maintaining a consistent view of the entire room for all participants. We achieve this by using two sets of cameras - a cluster of ”Panorama Cameras” located at the center of the display wall and are used to capture a panoramic view of the entire room, and a set of ”Personal Cameras” distributed along the display wall to capture front views of nearby participants. A robust segmentation algorithm with the assistance of depth cameras and an image synthesis algorithm work together to generate a consistent view of the entire scene. In our experience this new approach generates fewer distracting artifacts than conventional 3D reconstruction methods, while effectively correcting for eye gaze.
Mingsong Dou, Jan-Michael Frahm, Henry Fuchs, Bill Mauchly, Mod Marathe
VR3
2012 Special issue on Virtual Representations and Modeling of Large-scale environments (VRML)
Jan-Michael Frahm, Marc Pollefeys, Frank Dellaert, Jana Kosecka
Comput. Vis. Image Underst.1
2012 Hysteroscopy video summarization and browsing by estimating the physician's attention on video segments
Wilson Gavião, Jacob Scharcanski, Jan-Michael Frahm, Marc Pollefeys
Medical Image Anal.3
2011 Adaptive Scale Selection for Hierarchical Stereo
abstract
Hierarchical stereo provides an efficient coarse-to-fine mechanism for disparity map estimation.However, common drawbacks of such an approach include the loss of high frequency structures not observable at coarse scale levels, as well as the unrecoverable propagation of erroneous disparity estimates through the scale space.This paper presents an adaptive scale selection mechanism to determine a suitable resolution level from which to begin the hierarchical depth estimation process for each pixel.The proposed scale selection mechanism allows us to robustly implement variable cost aggregation in order to reduce the variability of the photo-consistency measure across scale space.We also incorporate a weighted shiftable window mechanism to enable error correction during coarse-to-fine depth refinement.Experiments illustrate the effectiveness of our approach in terms of disparity accuracy, while attaining a computational efficiency compromise between full resolution and hierarchical disparity map estimation.
Yi-Hung Jen, Enrique Dunn, Pierre Fite Georgel, Jan-Michael Frahm
BMVC4
2011 iSpy: automatic reconstruction of typed input from compromising reflections
abstract
We investigate the implications of the ubiquity of personal mobile devices and reveal new techniques for compromising the privacy of users typing on virtual keyboards. Specifi- cally, we show that so-called compromising reflections (in, for example, a victim's sunglasses) of a device's screen are sufficient to enable automated reconstruction, from video, of text typed on a virtual keyboard. Despite our deliberate use of low cost commodity video cameras, we are able to compensate for variables such as arbitrary camera and device positioning and motion through the application of advanced computer vision and machine learning techniques. Using footage captured in realistic environments (e.g., on a bus), we show that we are able to reconstruct fluent translations of recorded data in almost all of the test cases, correcting users' typing mistakes at the same time. We believe these results highlight the importance of adjusting privacy expectations in response to emerging technologies.
Rahul Raguram, Andrew M. White 0002, Dibyendusekhar Goswami, Fabian Monrose, Jan-Michael Frahm
CCS5
2011 Online environment mapping
abstract
The paper proposes a vision based online mapping of large-scale environments. Our novel approach uses a hybrid representation of a fully metric Euclidean environment map and a topological map. This novel hybrid representation facilitates our scalable online hierarchical bundle adjustment approach. The proposed method achieves scalability by solving the local registration through embedding neighboring keyframes and landmarks into a Euclidean space. The global adjustment is performed on a segmentation of the keyframes and posed as the iterative optimization of the arrangement of keyframes in each segment and the arrangement of rigidly moving segments. The iterative global adjustment is performed concurrently with the local registration of the keyframes in a local map. Thus the map is always locally metric around the current location, and likely to be globally consistent. Loop closures are handled very efficiently benefiting from the topological nature of the map and overcoming the loss of the metric map properties as previous approaches. The effectiveness of the proposed method is demonstrated in real-time on various challenging video sequences.
Jongwoo Lim, Jan-Michael Frahm, Marc Pollefeys
CVPR2
2011 Repetition-based dense single-view reconstruction
abstract
This paper presents a novel approach for dense reconstruction from a single-view of a repetitive scene structure. Given an image and its detected repetition regions, we model the shape recovery as the dense pixel correspondences within a single image. The correspondences are represented by an interval map that tells the distance of each pixel to its matched pixels within the single image. In order to obtain dense repetitive structures, we develop a new repetition constraint that penalizes the inconsistency between the repetition intervals of the dynamically corresponding pixel pairs. We deploy a graph-cut to balance between the high-level constraint of geometric repetition and the low-level constraints of photometric consistency and spatial smoothness. We demonstrate the accurate reconstruction of dense 3D repetitive structures through a variety of experiments, which prove the robustness of our approach to outliers such as structure variations, illumination changes, and occlusions.
Changchang Wu, Jan-Michael Frahm, Marc Pollefeys
CVPR2
2011 A geometric solver for calibrated stereo egomotion
abstract
This paper introduces a novel geometrical solution for the pose estimation of a stereo camera system as commonly used in robotics, where the camera system balances between coverage and overlap. The proposed approach considers a set of features observed, respectively, in four, three and two views. In contrast to most algebraic solutions our constraints are geometrically meaningful. Initially, we use a four view feature to restrict our translation vector to lie on the surface of a sphere while setting orientation as a function of translation up to a single rotational degree of freedom. Next, we use a three view feature to restrict the translation vector to lie on a circle on the sphere, while completely defining orientation as a function of translation. Finally, we use a two view feature to determine the translation vector lying on the intersection of the circle and one of the generator lines of a doubly ruled quadric. We show how for this final step, the problem can be reduced to the intersection of two coplanar circles. We also analyze the degenerate configurations of the proposed solver and perform an experimental evaluation.
Enrique Dunn, Brian Clipp, Jan-Michael Frahm
ICCV3
2011 RECON: Scale-adaptive robust estimation via Residual Consensus
abstract
In this paper, we present a novel, threshold-free robust estimation framework capable of efficiently fitting models to contaminated data. While RANSAC and its many variants have emerged as popular tools for robust estimation, their performance is largely dependent on the availability of a reasonable prior estimate of the inlier threshold. In this work, we aim to remove this threshold dependency. We build on the observation that models generated from uncontaminated minimal subsets are “consistent” in terms of the behavior of their residuals, while contaminated models exhibit uncorrelated behavior. By leveraging this observation, we then develop a very simple, yet effective algorithm that does not require apriori knowledge of either the scale of the noise, or the fraction of uncontaminated points. The resulting estimator, RECON (REsidual CONsensus), is capable of elegantly adapting to the contamination level of the data, and shows excellent performance even at low inlier ratios and high noise levels. We demonstrate the efficiency of our framework on a variety of challenging estimation problems.
Rahul Raguram, Jan-Michael Frahm
ICCV2
2011 Modeling and Recognition of Landmark Image Collections Using Iconic Scene Graphs
Rahul Raguram, Changchang Wu, Jan-Michael Frahm, Svetlana Lazebnik
Int. J. Comput. Vis.3
2011 Maximum likelihood autocalibration
Stuart B. Heinrich, Wesley E. Snyder, Jan-Michael Frahm
Image Vis. Comput.3
2011 Feature tracking and matching in video using programmable graphics hardware
Sudipta N. Sinha, Jan-Michael Frahm, Marc Pollefeys, Yakup Genc
Mach. Vis. Appl.2
2010 Piecewise planar and non-planar stereo for urban scene reconstruction
abstract
Piecewise planar models for stereo have recently become popular for modeling indoor and urban outdoor scenes. The strong planarity assumption overcomes the challenges presented by poorly textured surfaces, and results in low complexity 3D models for rendering, storage, and transmission. However, such a model performs poorly in the presence of non-planar objects, for example, bushes, trees, and other clutter present in many scenes. We present a stereo method capable of handling more general scenes containing both planar and non-planar regions. Our proposed technique segments an image into piecewise planar regions as well as regions labeled as non-planar. The non-planar regions are modeled by the results of a standard multi-view stereo algorithm. The segmentation is driven by multi-view photoconsistency as well as the result of a color-and texture-based classifier, learned from hand-labeled planar and non-planar image regions. Additionally our method links and fuses plane hypotheses across multiple overlapping views, ensuring a consistent 3D reconstruction over an arbitrary number of images. Using our system, we have reconstructed thousands of frames of street-level video. Results show our method successfully recovers piecewise planar surfaces alongside general 3D surfaces in challenging scenes containing large buildings as well as residential houses.
David Gallup, Jan-Michael Frahm, Marc Pollefeys
CVPR2
2010 Building Rome on a Cloudless Day
Jan-Michael Frahm, Pierre Fite Georgel, David Gallup, Tim Johnson, Rahul Raguram, Changchang Wu, Yi-Hung Jen, Enrique Dunn, Brian Clipp, Svetlana Lazebnik
ECCV (4)1
2010 Detecting Large Repetitive Structures with Salient Boundaries
Changchang Wu, Jan-Michael Frahm, Marc Pollefeys
ECCV (2)2
2010 Parallel, real-time visual SLAM
abstract
In this paper we present a novel system for real-time, six degree of freedom visual simultaneous localization and mapping using a stereo camera as the only sensor. The system makes extensive use of parallelism both on the graphics processor and through multiple CPU threads. Working together these threads achieve real-time feature tracking, visual odometry, loop detection and global map correction using bundle adjustment. The resulting corrections are fed back into to the visual odometry system to limit its drift over long sequences. We demonstrate our system on a series videos from challenging indoor environments with moving occluders, visually homogenous regions with few features, scene parts with large changes in lighting and fast camera motion. The total system performs its task of global map building in real time including loop detection and bundle adjustment on typical office building scale scenes.
Brian Clipp, Jongwoo Lim, Jan-Michael Frahm, Marc Pollefeys
IROS3
2010 Joint radiometric calibration and feature tracking system with an application to stereo
Seon Joo Kim, David Gallup, Jan-Michael Frahm, Marc Pollefeys
Comput. Vis. Image Underst.3
2009 3D Motion Segmentation Using Intensity Trajectory
Greg Welch, Jan-Michael Frahm, Marc Pollefeys
ACCV (1)3
2009 Next Best View Planning for Active Model Improvement
abstract
We propose a novel approach to determining the Next Best View (NBV) for the task of efficiently building highly accurate 3D models from images. Our proposed method deploys a hierarchical uncertainty driven model refinement process designed to select vantage viewpoints based on the model’s covariance structure and appearance, as well as the camera characteristics. The developed NBV planning system incrementally builds a sensing strategy by sequentially finding the single camera placement, which best reduces an existing model’s 3D uncertainty. The generic nature of our system’s design and internal data representation makes it well suited to be applied to a wide variety of 3D modeling algorithms. It can be used within active computer vision systems as well as for optimized view selection from the set of available views. Experimental results are presented to illustrate the effectiveness and versatility of our approach.
Enrique Dunn, Jan-Michael Frahm
BMVC2
2009 From structure-from-motion point clouds to fast location recognition
abstract
Efficient view registration with respect to a given 3D reconstruction has many applications like inside-out tracking in indoor and outdoor environments, and geo-locating images from large photo collections. We present a fast location recognition technique based on structure from motion point clouds. Vocabulary tree-based indexing of features directly returns relevant fragments of 3D models instead of documents from the images database. Additionally, we propose a compressed 3D scene representation which improves recognition rates while simultaneously reducing the computation time and the memory consumption. The design of our method is based on algorithms that efficiently utilize modern graphics processing units to deliver real-time performance for view registration. We demonstrate the approach by matching hand-held outdoor videos to known 3D urban models, and by registering images from online photo collections to the corresponding landmarks.
Arnold Irschara, Christopher Zach, Jan-Michael Frahm, Horst Bischof
CVPR3
2009 Continuous maximal flows and Wulff shapes: Application to MRFs
abstract
Convex and continuous energy formulations for low level vision problems enable efficient search procedures for the corresponding globally optimal solutions. In this work we extend the well-established continuous, isotropic capacity-based maximal flow framework to the anisotropic setting. By using powerful results from convex analysis, a very simple and efficient minimization procedure is derived. Further, we show that many important properties carry over to the new anisotropic framework, e.g. globally optimal binary results can be achieved simply by thresholding the continuous solution. In addition, we unify the anisotropic continuous maximal flow approach with a recently proposed convex and continuous formulation for Markov random fields, thereby allowing more general smoothness priors to be incorporated. Dense stereo results are included to illustrate the capabilities of the proposed approach.
Christopher Zach, Marc Niethammer, Jan-Michael Frahm
CVPR3
2009 A new minimal solution to the relative pose of a calibrated stereo camera with small field of view overlap
abstract
In this paper we present a new minimal solver for the relative pose of a calibrated stereo camera (i.e. a pair of rigidly mounted cameras). Our method is based on the fact that a feature visible in all four images (two image pairs acquired at two points in time) constrains the relative pose of the second stereo camera to lie on a sphere around this feature, which has a known, triangulated position in the first stereo camera coordinate frame. This constraint leaves three degrees of freedom; two for the location of the second camera on the sphere, and the third for the rotation in the respective tangent plane. We use three 2D correspondences, in particular two correspondences from the left (or right) camera and one correspondence from the other camera, to solve for these three remaining degrees of freedom. This approach is amenable to stereo cameras having a small overlap in their views. We present an efficient solution for this novel relative pose problem, describe the incorporation of our proposed solver into the RANSAC framework, evaluate its performance given noise and outliers, and demonstrate its use in a real-time structure from motion system.
Brian Clipp, Christopher Zach, Jan-Michael Frahm, Marc Pollefeys
ICCV3
2009 Exploiting uncertainty in random sample consensus
abstract
In this work, we present a technique for robust estimation, which by explicitly incorporating the inherent uncertainty of the estimation procedure, results in a more efficient robust estimation algorithm. In addition, we build on recent work in randomized model verification, and use this to characterize the `non-randomness' of a solution. The combination of these two strategies results in a robust estimation procedure that provides a significant speed-up over existing RANSAC techniques, while requiring no prior information to guide the sampling process. In particular, our algorithm requires, on average, 3-10 times fewer samples than standard RANSAC, which is in close agreement with theoretical predictions. The efficiency of the algorithm is demonstrated on a selection of geometric estimation problems.
Rahul Raguram, Jan-Michael Frahm, Marc Pollefeys
ICCV2
2009 Developing visual sensing strategies through next best view planning
abstract
We propose an approach for acquiring geometric 3D models using cameras mounted on autonomous vehicles and robots. Our method uses structure from motion techniques from computer vision to obtain the geometric structure of the scene. To achieve an efficient goal-driven resource deployment, we develop an incremental approach, which alternates between an accuracy-driven next best view determination and recursive path planning. The next best view is determined by a novel cost function that quantifies the expected contribution of future viewing configurations. A sensing path for robot motion towards the next best view is then achieved by a cost-driven recursive search of intermediate viewing configurations. We discuss some of the properties of our view cost function in the context of an iterative view planning process and present experimental results on a synthetic environment.
Enrique Dunn, Jur P. van den Berg, Jan-Michael Frahm
IROS3
2009 Towards Large-Scale Visual Mapping and Localization
Marc Pollefeys, Jan-Michael Frahm, Friedrich Fraundorfer, Christopher Zach, Changchang Wu, Brian Clipp, David Gallup
ISRR2
2009 Adaptive, real-time visual simultaneous localization and mapping
abstract
In this paper we present a real-time simultaneous localization and mapping system which uses a stereo camera as its only input. We combine the benefits of KLT feature tracking, which include high speed and robustness to repetitive features, with wide baseline features, which allow for feature matching after large camera motions. Updating the map of feature locations and camera poses is considerably more expensive than performing KLT tracking. For this reason we use the optical flow measured by the KLT tracker to adaptively select key frames for which we do a full map and camera pose update. In this way we limit the processing to only ¿interesting¿ parts of the video sequence. Additionally, we maintain a consistent scene scale at low cost by using a GPU implementation of multi-camera scene flow, a generalization of KLT to the motion of image features in three dimensions. The system uses multiple sub-maps; scalable, bag of features recognition and geometric verification to recover from motion estimation failure or ¿kidnapping¿. This architecture allows the robot to grow the existing map online and in real time while storing all of the data necessary for an off-line optimization to complete loops. We demonstrate the robustness of our system in a challenging indoor environment that includes semi-reflective glass walls and people moving in the scene.
Brian Clipp, Christopher Zach, Jongwoo Lim, Jan-Michael Frahm, Marc Pollefeys
WACV4
2008 Variable baseline/resolution stereo
abstract
We present a novel multi-baseline, multi-resolution stereo method, which varies the baseline and resolution proportionally to depth to obtain a reconstruction in which the depth error is constant. This is in contrast to traditional stereo, in which the error grows quadratically with depth, which means that the accuracy in the near range far exceeds that of the far range. This accuracy in the near range is unnecessarily high and comes at significant computational cost. It is, however, non-trivial to reduce this without also reducing the accuracy in the far range. Many datasets, such as video captured from a moving camera, allow the baseline to be selected with significant flexibility. By selecting an appropriate baseline and resolution (realized using an image pyramid), our algorithm computes a depthmap which has these properties: 1) the depth accuracy is constant over the reconstructed volume, 2) the computational effort is spread evenly over the volume, 3) the angle of triangulation is held constant w.r.t. depth. Our approach achieves a given target accuracy with minimal computational effort, and is orders of magnitude faster than traditional stereo.
David Gallup, Jan-Michael Frahm, Philippos Mordohai, Marc Pollefeys
CVPR2
2008 Radiometric calibration with illumination change for outdoor scene analysis
abstract
The images of an outdoor scene collected over time are valuable in studying the scene appearance variation which can lead to novel applications and help enhance existing methods that were constrained to controlled environments. However, the images do not reflect the true appearance of the scene in many cases due to the radiometric properties of the camera : the radiometric response function and the changing exposure. We introduce a new algorithm to compute the radiometric response function and the exposure of images given a sequence of images of a static outdoor scene where the illumination is changing. We use groups of pixels with constant behaviors towards the illumination change for the response estimation and introduce a sinusoidal lighting variation model representing the daily motion of the sun to compute the exposures.
Seon Joo Kim, Jan-Michael Frahm, Marc Pollefeys
CVPR2
2008 Simple calibration of non-overlapping cameras with a mirror
abstract
Calibrating a network of cameras with non-overlapping views is an important and challenging problem in computer vision. In this paper, we present a novel technique for camera calibration using a planar mirror. We overcome the need for all cameras to see a common calibration object directly by allowing them to see it through a mirror. We use the fact that the mirrored views generate a family of mirrored camera poses that uniquely describe the real camera pose. Our method consists of the following two steps: (1) using standard calibration methods to find the internal and external parameters of a set of mirrored camera poses, (2) estimating the external parameters of the real cameras from their mirrored poses by formulating constraints between them. We demonstrate our method on real and synthetic data for camera clusters with small overlap between the views and non-overlapping views.
Ram Krishan Kumar, Adrian Ilie, Jan-Michael Frahm, Marc Pollefeys
CVPR3
2008 3D model matching with Viewpoint-Invariant Patches (VIP)
abstract
The robust alignment of images and scenes seen from widely different viewpoints is an important challenge for camera and scene reconstruction. This paper introduces a novel class of viewpoint independent local features for robust registration and novel algorithms to use the rich information of the new features for 3D scene alignment and large scale scene reconstruction. The key point of our approach consists of leveraging local shape information for the extraction of an invariant feature descriptor. The advantages of the novel viewpoint invariant patch (VIP) are: that the novel features are invariant to 3D camera motion and that a single VIP correspondence uniquely defines the 3D similarity transformation between two scenes. In the paper we demonstrate how to use the properties of the VIPs in an efficient matching scheme for 3D scene alignment. The algorithm is based on a hierarchical matching method which tests the components of the similarity transformation sequentially to allow efficient matching and 3D scene alignment. We evaluate the novel features on real data with known ground truth information and show that the features can be used to reconstruct large scale urban scenes.
Changchang Wu, Brian Clipp, Xiaowei Li 0007, Jan-Michael Frahm, Marc Pollefeys
CVPR4
2008 Modeling and Recognition of Landmark Image Collections Using Iconic Scene Graphs
Xiaowei Li 0007, Changchang Wu, Christopher Zach, Svetlana Lazebnik, Jan-Michael Frahm
ECCV (1)5
2008 A Comparative Analysis of RANSAC Techniques Leading to Adaptive Real-Time Random Sample Consensus
Rahul Raguram, Jan-Michael Frahm, Marc Pollefeys
ECCV (2)2
2008 Robust 6DOF Motion Estimation for Non-Overlapping, Multi-Camera Systems
abstract
This paper introduces a novel, robust approach for 6DOF motion estimation of a multi-camera system with non-overlapping views. The proposed approach is able to solve the pose estimation, including scale, for a two camera system with non-overlapping views. In contrast to previous approaches, it degrades gracefully if the motion is close to degenerate. For degenerate motions the technique estimates the remaining 5DOF. The proposed technique is evaluated on real and synthetic sequences.
Brian Clipp, Jae-Hak Kim, Jan-Michael Frahm, Marc Pollefeys, Richard I. Hartley
WACV3
2008 Detailed Real-Time Urban 3D Reconstruction from Video
Marc Pollefeys, David Nistér, Jan-Michael Frahm, Amir Akbarzadeh, Philippos Mordohai, Brian Clipp, Chris Engels, David Gallup, Seon Joo Kim, Paul Merrell, C. Salmi, Sudipta N. Sinha, B. Talton, Liang Wang 0002, Qingxiong Yang, Henrik Stewénius, Ruigang Yang, Greg Welch, Herman Towles
Int. J. Comput. Vis.3
2007 Visual Odometry for Non-overlapping Views Using Second-Order Cone Programming
Jae-Hak Kim, Richard I. Hartley, Jan-Michael Frahm, Marc Pollefeys
ACCV (2)3
2007 Structure from Motion via Two-State Pipeline of Extended Kalman Filters
abstract
We introduce a novel approach to on-line structure from motion, using a pipelined pair of extended Kalman filters to improve accuracy with a minimal increase in computational cost. The two filters, a leading and a following filter, run concurrently on the same measurements in a synchronized producer-consumer fashion, but offset from each other in time. The leading filter estimates structure and motion using all of the available measurements from an optical flow based 2D tracker, passing the best 3D feature estimates, covariances, and associated measurements to the following filter, which runs several steps behind. This pipelined arrangement introduces a degree of noncausal behavior, effectively giving the following filter the benefit of decisions and estimates made several steps ahead. This means that the following filter works with only the best features, and can begin full 3D estimation from the very start of the respective 2D tracks. We demonstrate a reduction of more than 50% in mean reprojection errors using this approach on real data.
Brian Clipp, Greg Welch, Jan-Michael Frahm, Marc Pollefeys
BMVC3
2007 Real-Time Plane-Sweeping Stereo with Multiple Sweeping Directions
abstract
Recent research has focused on systems for obtaining automatic 3D reconstructions of urban environments from video acquired at street level. These systems record enormous amounts of video; therefore a key component is a stereo matcher which can process this data at speeds comparable to the recording frame rate. Furthermore, urban environments are unique in that they exhibit mostly planar surfaces. These surfaces, which are often imaged at oblique angles, pose a challenge for many window-based stereo matchers which suffer in the presence of slanted surfaces. We present a multi-view plane-sweep-based stereo algorithm which correctly handles slanted surfaces and runs in real-time using the graphics processing unit (GPU). Our algorithm consists of (1) identifying the scene's principle plane orientations, (2) estimating depth by performing a plane-sweep for each direction, (3) combining the results of each sweep. The latter can optionally be performed using graph cuts. Additionally, by incorporating priors on the locations of planes in the scene, we can increase the quality of the reconstruction and reduce computation time, especially for uniform textureless surfaces. We demonstrate our algorithm on a variety of scenes and show the improved accuracy obtained by accounting for slanted surfaces.
David Gallup, Jan-Michael Frahm, Philippos Mordohai, Qingxiong Yang, Marc Pollefeys
CVPR2
2007 Differential Camera Tracking through Linearizing the Local Appearance Manifold
abstract
The appearance of a scene is a function of the scene contents, the lighting, and the camera pose. A set of n-pixel images of a non-degenerate scene captured from different perspectives lie on a 6D nonlinear manifold in Rn. In general, this nonlinear manifold is complicated and numerous samples are required to learn it globally. In this paper, we present a novel method and some preliminary results for incrementally tracking camera motion through sampling and linearizing the local appearance manifold. At each frame time, we use a cluster of calibrated and synchronized small baseline cameras to capture scene appearance samples at different camera poses. We compute a first-order approximation of the appearance manifold around the current camera pose. Then, as new cluster samples are captured at the next frame time, we estimate the incremental camera motion using a linear solver. By using intensity measurements and directly sampling the appearance manifold, our method avoids the commonly-used feature extraction and matching processes, and does not require 3D correspondences across frames. Thus it can be used for scenes with complicated surface materials, geometries, and view-dependent appearance properties, situations where many other camera tracking methods would fail.
Marc Pollefeys, Greg Welch, Jan-Michael Frahm, Adrian Ilie
CVPR4
2007 Joint Feature Tracking and Radiometric Calibration from Auto-Exposure Video
abstract
To capture the full brightness range of natural scenes, cameras automatically adjust the exposure value which causes the brightness of scene points to change from frame to frame. Given such a video sequence, we introduce a new method for tracking features and estimating the radiometric response function of the camera and the exposure difference between frames simultaneously. We model the global and nonlinear process that is responsible for the changes in image brightness rather than adapting to the changes locally and linearly which makes our tracking more robust to the change in brightness. The radiometric response function and the exposure difference between frames are also estimated in the process which enables further video processing algorithms to deal with the varying brightness.
Seon Joo Kim, Jan-Michael Frahm, Marc Pollefeys
ICCV2
2007 Real-Time Visibility-Based Fusion of Depth Maps
abstract
We present a viewpoint-based approach for the quick fusion of multiple stereo depth maps. Our method selects depth estimates for each pixel that minimize violations of visibility constraints and thus remove errors and inconsistencies from the depth maps to produce a consistent surface. We advocate a two-stage process in which the first stage generates potentially noisy, overlapping depth maps from a set of calibrated images and the second stage fuses these depth maps to obtain an integrated surface with higher accuracy, suppressed noise, and reduced redundancy. We show that by dividing the processing into two stages we are able to achieve a very high throughput because we are able to use a computationally cheap stereo algorithm and because this architecture is amenable to hardware-accelerated (GPU) implementations. A rigorous formulation based on the notion of stability of a depth estimate is presented first. It aims to determine the validity of a depth estimate by rendering multiple depth maps into the reference view as well as rendering the reference depth map into the other views in order to detect occlusions and free- space violations. We also present an approximate alternative formulation that selects and validates only one hypothesis based on confidence. Both formulations enable us to perform video-based reconstruction at up to 25 frames per second. We show results on the multi-view stereo evaluation benchmark datasets and several outdoors video sequences. Extensive quantitative analysis is performed using an accurately surveyed model of a real building as ground truth.
Paul Merrell, Amir Akbarzadeh, Liang Wang 0002, Philippos Mordohai, Jan-Michael Frahm, Ruigang Yang, David Nistér, Marc Pollefeys
ICCV5
2007 Evaluation of Large Scale Scene Reconstruction
abstract
We present an evaluation methodology and data for large scale video-based 3D reconstruction. We evaluate the effects of several parameters and draw conclusions that can be useful for practical systems operating in uncontrolled environments. Unlike the benchmark datasets used for the binocular stereo and multi-view reconstruction evaluations, which were collected under well-controlled conditions, our datasets are captured outdoors using video cameras mounted on a moving vehicle. As a result, the videos are much more realistic and include phenomena such as exposure changes from viewing both bright and dim surfaces, objects at varying distances from the camera, and objects of varying size and degrees of texture. The dataset includes ground truth models and precise camera pose information. We also present an evaluation methodology applicable to reconstructions of large scale environments. We evaluate the accuracy and completeness of reconstructions obtained by two fast, visibility-based depth map fusion algorithms as parameters vary.
Paul Merrell, Philippos Mordohai, Jan-Michael Frahm, Marc Pollefeys
ICCV3
2006 RANSAC for (Quasi-)Degenerate data (QDEGSAC)
abstract
The computation of relations from a number of potential matches is a major task in computer vision. Often RANSAC is employed for the robust computation of relations such as the fundamental matrix. For (quasi-)degenerate data however, it often fails to compute the correct relation. The computed relation is always consistent with the data but RANSAC does not verify that it is unique. The paper proposes a framework that estimates the correct relation with the same robustness as RANSAC even for (quasi-)degenerate data. The approach is based on a hierarchical RANSAC over the number of constraints provided by the data. In contrast to all previously presented algorithms for (quasi-)degenerate data our technique does not require problem specific tests or models to deal with degenerate configurations. Accordingly it can be applied for the estimation of any relation on any data and is not limited to a special type of relation as previous approaches. The results are equivalent to the results achieved by state of the art approaches that employ knowledge about degeneracies.
Jan-Michael Frahm, Marc Pollefeys
CVPR (1)1
2006 Tutorials - MAR Tutorial 1 (half day) & ISMAR Tutorial 2 (half day)
abstract
Provides an abstract for each of the presentations and a brief professional biography of each presenter. The complete presentations were not made available for publication as part of the conference proceedings.
Jan-Michael Frahm, Jannick P. Rolland, Andrei State, Ozan Cakmakci
ISMAR1
2003 Camera Calibration with Known Rotation
abstract
We address the problem of using external rotation information with uncalibrated video sequences. The main problem addressed is, what is the benefit of the orientation information for camera calibration? It is shown that in case of a rotating camera the camera calibration problem is linear even in the case that all intrinsic parameters vary. For arbitrarily moving cameras the calibration problem is also linear but underdetermined for the general case of varying all intrinsic parameters. However, if certain constraints are applied to the intrinsic parameters the camera calibration can be computed linearly. It is analyzed which constraints are needed for camera calibration of freely moving cameras. Furthermore we address the problem of aligning the camera data with the rotation sensor data in time. We give an approach to align these data in case of a rotating camera.
Jan-Michael Frahm, Reinhard Koch
ICCV1