VLDB 2026 Research / reviewers in the wild / expert
Chris Sweeney
dblp:124/0715 · also Christopher Sweeney
· DBLP profile ↗
28ranked-venue papers
10as first author
7since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 8 first-author · 6 since 2021Artificial intelligence and machine learning · 16 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 first-authorSystems, architecture and hardware · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | EgoLifter: Open-World 3D Segmentation for Egocentric Perception
Qiao Gu, Zhaoyang Lv, Duncan P. Frost, Simon Green, Julian Straub, Chris Sweeney |
ECCV (43) | 6 |
| 2024 | STT: Stateful Tracking with Transformers for Autonomous DrivingabstractTracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model’s performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTPSthat address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset. Longlong Jing, Ruichi Yu, Zhengli Zhao, Shiwei Sheng, Colin Graber, Qinru Li, Shangxuan Wu, Chris Sweeney, Wei-Chih Hung, Xingyi Zhou, Farshid Moussavi, James Guo, Mingxing Tan, Weilong Yang |
ICRA | 12 |
| 2022 | LISA: Learning Implicit Shape and Appearance of HandsabstractThis paper proposes a do-it-all neural model of human hands, named LISA. The model can capture accurate hand shape and appearance, generalize to arbitrary hand sub-jects, provide dense surface correspondences, be reconstructed from images in the wild, and can be easily an-imated. We train LISA by minimizing the shape and appearance losses on a large set of multi-view RGB image se-quences annotated with coarse 3D poses of the hand skele-ton. For a 3D point in the local hand coordinates, our model predicts the color and the signed distance with respect to each hand bone independently, and then combines the per-bone predictions using the predicted skinning weights. The shape, color, and pose representations are disentangled by design, enabling fine control of the selected hand param-eters. We experimentally demonstrate that LISA can ac-curately reconstruct a dynamic hand from monocular or multi-view sequences, achieving a noticeably higher qual-ity of reconstructed hand shapes compared to baseline approaches. Project page: https://www.iri.upc.edu/people/ecorona/lisa/. Enric Corona, Tomas Hodan, Minh Vo, Francesc Moreno-Noguer, Chris Sweeney, Richard A. Newcombe, Lingni Ma |
CVPR | 5 |
| 2022 | NinjaDesc: Content-Concealing Visual Descriptors via Adversarial LearningabstractIn the light of recent analyses on privacy-concerning scene revelation from visual descriptors, we develop descriptors that conceal the input image content. In particular, we propose an adversarial learning framework for training visual descriptors that prevent image reconstruction, while maintaining the matching accuracy. We let a feature encoding network and image reconstruction network compete with each other, such that the feature encoder tries to impede the image reconstruction with its generated descriptors, while the reconstructor tries to recover the input image from the descriptors. The experimental results demonstrate that the visual descriptors obtained with our method significantly deteriorate the image reconstruction quality with minimal impact on correspondence matching and camera localization performance. Tony Ng, Hyo Jin Kim 0004, Vincent T. Lee, Daniel DeTone, Tsun-Yi Yang, Tianwei Shen, Eddy Ilg, Vassileios Balntas, Krystian Mikolajczyk, Chris Sweeney |
CVPR | 10 |
| 2022 | Self-supervised Neural Articulated Shape and Appearance ModelsabstractLearning geometry, motion, and appearance priors of object classes is important for the solution of a large variety of computer vision problems. While the majority of approaches has focused on static objects, dynamic objects, especially with controllable articulation, are less explored. We propose a novel approach for learning a representation of the geometry, appearance, and motion of a class of articulated objects given only a set of color images as input. In a self-supervised manner, our novel representation learns shape, appearance, and articulation codes that enable independent control of these semantic dimensions. Our model is trained end-to-end without requiring any articulation annotations. Experiments show that our approach performs well for different joint types, such as revolute and prismatic joints, as well as different combinations of these joints. Compared to state of the art that uses direct 3D supervision and does not output appearance, we recover more faithful geometry and appearance from 2D observations only. In addition, our representation enables a large variety of applications, such as few-shot reconstruction, the generation of novel articulations, and novel view-synthesis. Project page: https://weify627.github.io/nasam/. Fangyin Wei, Rohan Chabra, Lingni Ma, Christoph Lassner, Michael Zollhöfer, Szymon Rusinkiewicz, Chris Sweeney, Richard A. Newcombe, Mira Slavcheva |
CVPR | 7 |
| 2022 | Relative Pose Solvers using Monocular DepthabstractWe describe a novel approach for using deep-learned priors to estimate the pose of a camera and show that these priors can be efficiently and accurately used for robust relative pose estimation. We use an off-the-shelf monocular depth network to provide an estimation of up-to-scale depth per pixel, and propose three new methods for solving for relative pose as well as a new algorithm for homography estimation. The additional signal provided by the depths leads to efficient solvers that require fewer correspondences than traditional methods and provide accurate and robust pose estimation when combined with state-of-the-art robust estimators, e.g., Graph-Cut RANSAC. The algorithms are tested on more than 70,000 publicly available image pairs from the 1DSfM dataset. The accuracy of the proposed methods are comparable or better than the standard five-point algorithm, and the reduced number of necessary correspondences speed up the robust estimation procedure, sometimes by orders of magnitude. Daniel Barath, Chris Sweeney |
ICPR | 2 |
| 2021 | ODAM: Object Detection, Association, and Mapping using Posed RGB VideoabstractLocalizing objects and estimating their extent in 3D is an important step towards high-level 3D scene understanding, which has many applications in Augmented Reality and Robotics. We present ODAM, a system for 3D Object Detection, Association, and Mapping using posed RGB videos. The proposed system relies on a deep learning front-end to detect 3D objects from a given RGB frame and associate them to a global object-based map using a graph neural network (GNN). Based on these frame-to-model associations, our back-end optimizes object bounding volumes, represented as super-quadrics, under multi-view geometry constraints and the object scale prior. We validate the proposed system on ScanNet where we show a significant improvement over existing RGB-only methods. Kejie Li, Daniel DeTone, Steven Chen, Minh Vo, Ian D. Reid 0001, Seyed Hamid Rezatofighi, Chris Sweeney, Julian Straub, Richard A. Newcombe |
ICCV | 7 |
| 2020 | Reducing Drift in Structure From Motion Using Extended FeaturesabstractLow-frequency long-range errors (drift) are an endemic problem in 3D structure from motion, and can often hamper reasonable reconstructions of the scene. In this paper, we present a method to dramatically reduce scale and positional drift by using extended structural features such as planes and vanishing points. Unlike traditional feature matches, our extended features are able to span non-overlapping input images, and hence provide long-range constraints on the scale and shape of the reconstruction. We add these features as additional constraints to a state-of the-art global structure from motion algorithm and demonstrate that the added constraints enable the reconstruction of particularly drift-prone sequences such as long, low field-of-view videos without inertial measurements. Additionally, we provide an analysis of the drift-reducing capabilities of these constraints by evaluating on a synthetic dataset. Our structural features are able to significantly reduce drift for scenes that contain long-spanning man-made structures, such as aligned rows of windows or planar building facades. Aleksander Holynski, David Geraghty, Jan-Michael Frahm, Chris Sweeney, Richard Szeliski |
3DV | 4 |
| 2020 | Domain Adaptation of Learned Featuresfor Visual Localization
Sungyong Baik, Hyo Jin Kim 0004, Tianwei Shen, Eddy Ilg, Kyoung Mu Lee, Chris Sweeney |
BMVC | 6 |
| 2019 | A Transparent Framework for Evaluating Unintended Demographic Bias in Word EmbeddingsabstractWord embedding models have gained a lot of traction in the Natural Language Processing community, however, they suffer from unintended demographic biases. Most approaches to evaluate these biases rely on vector space based metrics like the Word Embedding Association Test (WEAT). While these approaches offer great geometric insights into unintended biases in the embedding vector space, they fail to offer an interpretable meaning for how the embeddings could cause discrimination in downstream NLP applications. In this work, we present a transparent framework and metric for evaluating discrimination across protected groups with respect to their word embedding bias. Our metric (Relative Negative Sentiment Bias, RNSB) measures fairness in word embeddings via the relative negative sentiment associated with demographic identity terms from various protected groups. We show that our framework and metric enable useful analysis into the bias in word embeddings. Chris Sweeney, Maryam Najafian |
ACL (1) | 1 |
| 2019 | StereoDRNet: Dilated Residual StereoNetabstractWe propose a system that uses a convolution neural network (CNN) to estimate depth from a stereo pair followed by volumetric fusion of the predicted depth maps to produce a 3D reconstruction of a scene. Our proposed depth refinement architecture, predicts view-consistent disparity and occlusion maps that helps the fusion system to produce geometrically consistent reconstructions. We utilize 3D dilated convolutions in our proposed cost filtering network that yields better filtering while almost halving the computational cost in comparison to state of the art cost filtering architectures. For feature extraction we use the Vortex Pooling architecture. The proposed method achieves state of the art results in KITTI 2012, KITTI 2015 and ETH 3D stereo benchmarks. Finally, we demonstrate that our system is able to produce high fidelity 3D scene reconstructions that outperforms the state of the art stereo system. Rohan Chabra, Julian Straub, Chris Sweeney, Richard A. Newcombe, Henry Fuchs |
CVPR | 3 |
| 2019 | A Supervised Approach to Predicting Noise in Depth ImagesabstractModern robotic systems are very complex and need to be tested in simulations with detailed sensor noise models to effectively verify robotic behavior. Depth imagery in particular comes with significant noise in the form of scene-dependent pixel-wise dropouts and distortions. Unfortunately, many depth camera simulations contain limited noise models, or can only support generating realistic depth images of simple scenes, which limits their usefulness in effectively testing perception algorithms. We propose a data driven approach to generate more realistic noise for complex simulated environments by using a convolutional neural network (CNN) to predict which pixels of a simulated noise-free depth image will not have returns (no-depth-return pixels, or NDP). We choose to focus on NDP here, as these dropouts are the most common and dramatic form of depth image noise. To train this network, we use reconstructed real-world scenes from the Label Fusion dataset to provide ground truth depth for each noisy depth image used to scan the scene. We use the resulting noise-free and noisy depth image pairs as labeled examples and train the network to predict which pixels of the noise-free image will be NDP. When used to post-process a simulation of a depth sensor, this system produces realistic depth images, even in cluttered scenes. To demonstrate that our approach successfully closes the reality gap for depth imagery, we show that the popular ICP algorithm for object pose estimation fails more realistically on our CNN-corrupted simulated depth images than on uncorrupted depth images and unsupervised domain adaptation baselines. Chris Sweeney, Gregory Izatt, Russ Tedrake |
ICRA | 1 |
| 2017 | GraphMatch: Efficient Large-Scale Graph Construction for Structure from MotionabstractWe present GraphMatch, an approximate yet efficient method for building the matching graph for large-scale structure-from-motion~(SfM) pipelines. GraphMatch leverages two priors that can predict which image pairs are likely to match, thereby making the matching process for SfM much more efficient. The first is a score computed from the distance between the Fisher vectors of any two images. The second prior is based on the graph distance between vertices in the underlying matching graph. GraphMatch combines these two priors into an iterative ``sample-and-propagate'' scheme similar to the PatchMatch algorithm. Its sampling stage uses Fisher similarity priors to guide the search for matching image pairs, while its propagation stage explores neighbors of matched pairs to find new ones with a high image similarity score. Our experiments show that GraphMatch finds the most image pairs as compared to competing, approximate methods while at the same time being the most efficient. Qiaodong Cui, Victor Fragoso, Chris Sweeney, Pradeep Sen |
3DV | 3 |
| 2017 | ANSAC: Adaptive Non-Minimal Sample and Consensus
Victor Fragoso, Chris Sweeney, Pradeep Sen, Matthew Turk 0001 |
BMVC | 2 |
| 2016 | Large Scale SfM with the Distributed Camera ModelabstractWe introduce the distributed camera model, a novel model for Structure-from-Motion (SfM). This model describes image observations in terms of light rays with ray origins and directions rather than pixels. As such, the proposed model is capable of describing a single camera or multiple cameras simultaneously as the collection of all light rays observed. We show how the distributed camera model is a generalization of the standard camera model and we describe a general formulation and solution to the absolute camera pose problem that works for standard or distributed cameras. The proposed method computes a solution that is up to 8 times more efficient and robust to rotation singularities in comparison with gDLS[21]. Finally, this method is used in an novel large-scale incremental SfM pipeline where distributed cameras are accurately and robustly merged together. This pipeline is a direct generalization of traditional incremental SfM, however, instead of incrementally adding one camera at a time to grow the reconstruction the reconstruction is grown by adding a distributed camera. Our pipeline produces highly accurate reconstructions efficiently by avoiding the need for many bundle adjustment iterations and is capable of computing a 3D model of Rome from over 15,000 images in just 22 minutes. Chris Sweeney, Victor Fragoso, Tobias Höllerer, Matthew Turk 0001 |
3DV | 1 |
| 2016 | Multi-view gesture annotations in image-based 3D reconstructed scenesabstractWe present a novel 2D gesture annotation method for use in image-based 3D reconstructed scenes with applications in collaborative virtual and augmented reality. Image-based reconstructions allow users to virtually explore a remote environment using image-based rendering techniques. To collaborate with other users, either synchronously or asynchronously, simple 2D gesture annotations can be used to convey spatial information to another user. Unfortunately, prior methods are either unable to disambiguate such 2D annotations in 3D from novel viewpoints or require relatively dense reconstructions of the environment. Benjamin Nuernberger, Kuo-Chin Lien, Lennon Grinta, Chris Sweeney, Matthew Turk 0001, Tobias Höllerer |
VRST | 4 |
| 2016 | The generalized relative pose and scale problem: View-graph fusion via 2D-2D registrationabstractIt is well-known that the relative pose problem can be generalized to non-central cameras. We present a further generalization, denoted the generalized relative pose and scale problem. It has surprising importance for classical problems such as solving similarity transformations for view-graph concatenation in hierarchical structure from motion and loop-closure in visual SLAM, both posed as a 2D-2D registration problem. The relative pose problem and all its generalizations constitute a family of similar symmetric eigenvalue problems, which allow us to compress data and find a geometrically meaningful solution by an efficient search in the space of rotations. While the derivation of a completely general closed-form solver appears intractable, we make use of a simple heuristic global energy minimization scheme based on local minimum suppression, returning outstanding performance in practically relevant scenarios. Efficiency and reliability of our algorithm are demonstrated on both simulated and real data, supporting our claim of superior performance with respect to both generalized 2D-3D and 3D-3D registration approaches. By directly employing image information, we avoid the common noise in point clouds occurring especially along the depth direction. Laurent Kneip, Chris Sweeney, Richard I. Hartley |
WACV | 2 |
| 2016 | Computational Reconstruction of NFκB Pathway Interaction Mechanisms during Prostate CancerabstractMolecular research in cancer is one of the largest areas of bioinformatic investigation, but it remains a challenge to understand biomolecular mechanisms in cancer-related pathways from high-throughput genomic data. This includes the Nuclear-factor-kappa-B (NFκB) pathway, which is central to the inflammatory response and cell proliferation in prostate cancer development and progression. Despite close scrutiny and a deep understanding of many of its members' biomolecular activities, the current list of pathway members and a systems-level understanding of their interactions remains incomplete. Here, we provide the first steps toward computational reconstruction of interaction mechanisms of the NFκB pathway in prostate cancer. We identified novel roles for ATF3, CXCL2, DUSP5, JUNB, NEDD9, SELE, TRIB1, and ZFP36 in this pathway, in addition to new mechanistic interactions between these genes and 10 known NFκB pathway members. A newly predicted interaction between NEDD9 and ZFP36 in particular was validated by co-immunoprecipitation, as was NEDD9's potential biological role in prostate cancer cell growth regulation. We combined 651 gene expression datasets with 1.4M gene product interactions to predict the inclusion of 40 additional genes in the pathway. Molecular mechanisms of interaction among pathway members were inferred using recent advances in Bayesian data integration to simultaneously provide information specific to biological contexts and individual biomolecular activities, resulting in a total of 112 interactions in the fully reconstructed NFκB pathway: 13 (11%) previously known, 29 (26%) supported by existing literature, and 70 (63%) novel. This method is generalizable to other tissue types, cancers, and organisms, and this new information about the NFκB pathway will allow us to further understand prostate cancer and to develop more effective prevention and treatment strategies. Daniela Börnigen, Svitlana Tyekucheva, Jennifer R. Rider, Gwo-Shu Lee, Lorelei A. Mucci, Chris Sweeney, Curtis Huttenhower |
PLoS Comput. Biol. | 7 |
| 2015 | Computing similarity transformations from only image correspondencesabstractWe propose a novel solution for computing the relative pose between two generalized cameras that includes reconciling the internal scale of the generalized cameras. This approach can be used to compute a similarity transformation between two coordinate systems, making it useful for loop closure in visual odometry and registering multiple structure from motion reconstructions together. In contrast to alternative similarity transformation methods, our approach uses 2D-2D image correspondences thus is not subject to the depth uncertainty that often arises with 3D points. We utilize a known vertical direction (which may be easily obtained from IMU data or vertical vanishing point detection) of the generalized cameras to solve the generalized relative pose and scale problem as an efficient Quadratic Eigenvalue Problem. To our knowledge, this is the first method for computing similarity transformations that does not require any 3D information. Our experiments on synthetic and real data demonstrate that this leads to improved performance compared to methods that use 3D-3D or 2D-3D correspondences, especially as the depth of the scene increases. Chris Sweeney, Laurent Kneip, Tobias Höllerer, Matthew Turk 0001 |
CVPR | 1 |
| 2015 | Optimizing the Viewing Graph for Structure-from-MotionabstractThe viewing graph represents a set of views that are related by pairwise relative geometries. In the context of Structure-from-Motion (SfM), the viewing graph is the input to the incremental or global estimation pipeline. Much effort has been put towards developing robust algorithms to overcome potentially inaccurate relative geometries in the viewing graph during SfM. In this paper, we take a fundamentally different approach to SfM and instead focus on improving the quality of the viewing graph before applying SfM. Our main contribution is a novel optimization that improves the quality of the relative geometries in the viewing graph by enforcing loop consistency constraints with the epipolar point transfer. We show that this optimization greatly improves the accuracy of relative poses in the viewing graph and removes the need for filtering steps or robust algorithms typically used in global SfM methods. In addition, the optimized viewing graph can be used to efficiently calibrate cameras at scale. We combine our viewing graph optimization and focal length calibration into a global SfM pipeline that is more efficient than existing approaches. To our knowledge, ours is the first global SfM pipeline capable of handling uncalibrated image sets. Chris Sweeney, Torsten Sattler, Tobias Höllerer, Matthew Turk 0001, Marc Pollefeys |
ICCV | 1 |
| 2015 | Efficient Computation of Absolute Pose for Gravity-Aware Augmented RealityabstractWe propose a novel formulation for determining the absolute pose of a single or multi-camera system given a known vertical direction. The vertical direction may be easily obtained by detecting the vertical vanishing points with computer vision techniques, or with the aid of IMU sensor measurements from a smartphone. Our solver is general and able to compute absolute camera pose from two 2D-3D correspondences for single or multi-camera systems. We run several synthetic experiments that demonstrate our algorithm's improved robustness to image and IMU noise compared to the current state of the art. Additionally, we run an image localization experiment that demonstrates the accuracy of our algorithm in real-world scenarios. Finally, we show that our algorithm provides increased performance for real-time model-based tracking compared to solvers that do not utilize the vertical direction and show our algorithm in use with an augmented reality application running on a Google Tango tablet. Chris Sweeney, John Flynn, Benjamin Nuernberger, Matthew Turk 0001, Tobias Höllerer |
ISMAR | 1 |
| 2015 | Theia: A Fast and Scalable Structure-from-Motion LibraryabstractIn this paper, we have presented a comprehensive multi-view geometry library, Theia, that focuses on large-scale SfM. In addition to state-of-the-art scalable SfM pipelines, the library provides numerous tools that are useful for students, researchers, and industry experts in the field of multi-view geometry. Theia contains clean code that is well documented (with code comments and the website) and easy to extend. The modular design allows for users to easily implement and experiment with new algorithms within our current pipeline without having to implement a full end-to-end SfM pipeline themselves. Theia has already gathered a large number of diverse users from universities, startups, and industry and we hope to continue to gather users and active contributors from the open-source community. Chris Sweeney, Tobias Höllerer, Matthew Turk 0001 |
ACM Multimedia | 1 |
| 2014 | Solving for Relative Pose with a Partially Known Rotation is a Quadratic Eigenvalue ProblemabstractWe propose a novel formulation of minimal case solutions for determining the relative pose of perspective and generalized cameras given a partially known rotation, namely, a known axis of rotation. An axis of rotation may be easily obtained by detecting vertical vanishing points with computer vision techniques, or with the aid of sensor measurements from a smart phone. Given a known axis of rotation, our algorithms solve for the angle of rotation around the known axis along with the unknown translation. We formulate these relative pose problems as Quadratic Eigen value Problems which are very simple to construct. We run several experiments on synthetic and real data to compare our methods to the current state-of-the-art algorithms. Our methods provide several advantages over alternatives methods, including efficiency and accuracy, particularly in the presence of image and sensor noise as is often the case for mobile devices. Chris Sweeney, John Flynn, Matthew Turk 0001 |
3DV | 1 |
| 2014 | On Sampling Focal Length Values to Solve the Absolute Pose Problem
Torsten Sattler, Chris Sweeney, Marc Pollefeys |
ECCV (4) | 2 |
| 2014 | gDLS: A Scalable Solution to the Generalized Pose and Scale Problem
Chris Sweeney, Victor Fragoso, Tobias Höllerer, Matthew Turk 0001 |
ECCV (4) | 1 |
| 2014 | Model Estimation and Selection towardsUnconstrained Real-Time Tracking and MappingabstractWe present an approach and prototype implementation to initialization-free real-time tracking and mapping that supports any type of camera motion in 3D environments, that is, parallax-inducing as well as rotation-only motions. Our approach effectively behaves like a keyframe-based Simultaneous Localization and Mapping system or a panorama tracking and mapping system, depending on the camera movement. It seamlessly switches between the two modes and is thus able to track and map through arbitrary sequences of parallax-inducing and rotation-only camera movements. The system integrates both model-based and model-free tracking, automatically choosing between the two depending on the situation, and subsequently uses the "Geometric Robust Information Criterion" to decide whether the current camera motion can best be represented as a parallax-inducing motion or a rotation-only motion. It continues to collect and map data after tracking failure by creating separate tracks which are later merged if they are found to overlap. This is in contrast to most existing tracking and mapping systems, which suspend tracking and mapping and thus discard valuable data until relocalization with respect to the initial map is successful. We tested our prototype implementation on a variety of video sequences, successfully tracking through different camera motions and fully automatically building combinations of panoramas and 3D structure. Steffen Gauglitz, Chris Sweeney, Jonathan Ventura, Matthew Turk 0001, Tobias Höllerer |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2013 | Improved outdoor augmented reality through "Globalization"abstractDespite the major interest in live tracking and mapping (e.g., SLAM), the field of augmented reality has yet to truly make use of the rich data provided from large-scale reconstructions generated by structure from motion. This dissertation focuses on extensible tracking and mapping for large-scale reconstructions that enables SfM and SLAM to operate cooperatively to mutually enhance the performance. We describe a multi-user, collaborative augmented reality system that will collectively extend and enhance reconstructions of urban environments at city-scales. Contrary to current outdoor augmented reality systems, this system is capable of continuous tracking through areas previously modeled as well as new, undiscovered areas. Further, we describe a new process called globalization that propagates new visual information back to the global model. Globalization allows for continuous updating of the 3D models with visual data from live users, providing data to fill coverage gaps that are common in 3D reconstructions and to provide the most current view of an environment as it changes over time. The proposed research is a crucial step toward enabling users to augment urban environments with location-specific information at any location in the world for a truly global augmented reality. Chris Sweeney, Tobias Höllerer, Matthew Turk 0001 |
ISMAR | 1 |
| 2012 | Live tracking and mapping from both general and rotation-only camera motionabstractWe present an approach to real-time tracking and mapping that supports any type of camera motion in 3D environments, that is, general (parallax-inducing) as well as rotation-only (degenerate) motions. Our approach effectively generalizes both a panorama mapping and tracking system and a keyframe-based Simultaneous Localization and Mapping (SLAM) system, behaving like one or the other depending on the camera movement. It seamlessly switches between the two and is thus able to track and map through arbitrary sequences of general and rotation-only camera movements. Key elements of our approach are to design each system component such that it is compatible with both panoramic data and Structure-from-Motion data, and the use of the `Geometric Robust Information Criterion' to decide whether the transformation between a given pair of frames can best be modeled with an essential matrix E, or with a homography H. Further key features are that no separate initialization step is needed, that the reconstruction is unbiased, and that the system continues to collect and map data after tracking failure, thus creating separate tracks which are later merged if they overlap. The latter is in contrast to most existing tracking and mapping systems, which suspend tracking and mapping, thus discarding valuable data, while trying to relocalize the camera with respect to the initial map. We tested our system on a variety of video sequences, successfully tracking through different camera motions and fully automatically building panoramas as well as 3D structures. Steffen Gauglitz, Chris Sweeney, Jonathan Ventura, Matthew Turk 0001, Tobias Höllerer |
ISMAR | 2 |