Clemens Arth

dblp:55/398 · DBLP profile ↗
← Back
33ranked-venue papers
8as first author
9since 2021 · last 2026
0000-0001-6949-4713ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 8 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 1 since 2021Systems, architecture and hardware · 2
YearPublicationVenuePosition
2026 Shape-Shifting Splats: Realtime Context Translation for Gaussian Splatting in VR
abstract
3D Gaussian Splatting (3DGS) has revolutionized novel-view synthesis by producing photorealistic renderings of virtually any object. However, 3DGS methods primarily support static scene representations, limiting flexibility and preventing direct application in scenarios that require visual variety. Our novel approach, Shape-Shifting Splats, addresses this limitation by enabling dynamic runtime transformations of Gaussian splats. Our versatile formulation allows changes in both color and geometry during render time, without increasing memory usage or introducing visual artifacts. Once trained, our Shape-Shifting Splats generate high-fidelity novel views while supporting modifications in geometry and appearance of the underlying model. In addition to applications in gaming and e-commerce, our method is applicable to mobile, industrial maintenance, and inspection. With multimodal data (e.g., thermal and RGB), we create interactive 3D representations that allow inspectors wearing VR glasses to seamlessly switch between thermal and RGB views, helping to prevent oversights.
Thomas Kernbauer, Simon Fussi, Philipp Fleck, Clemens Arth
VR4
2026 Change-Resilient Localization Estimation
abstract
Indoor localization is essential in applications such as augmented reality or robotics. Existing solutions for localization in static scenes work well even for large environments, but localization in environments with movable objects whose pose in the scene change between sessions remains challenging. In this paper, we propose a change-resilient localization method based on a novel geometric descriptor computed only from geometric primitives. Our method is capable of re-identifying primitives that have moved in the scene. We leverage this feature to update a stored reference model (anchor) of the environment to accommodate the changes, which enables localization that is resilient to changes in the scene. We report on a set of experiments demonstrating the robustness and scalability of our method. In addition, we present use cases highlighting the importance of being able to update a reference model.
Fernando Reyes-Aviles, Philipp Fleck, Dieter Schmalstieg, Clemens Arth
VR4
2026 Virtual memory for 3D Gaussian Splatting
abstract
3D Gaussian Splatting represents a breakthrough in the field of novel view synthesis. It establishes Gaussians as core rendering primitives for highly accurate real-world environment reconstruction. Recent advances have drastically increased the size of scenes that can be created. In this work, we present a method for rendering large and complex 3D Gaussian Splatting scenes using virtual memory. By leveraging well-established virtual memory and virtual texturing techniques, our approach efficiently identifies visible Gaussians and dynamically streams them to the GPU just in time for real-time rendering. Selecting only the necessary Gaussians for both storage and rendering results in reduced memory usage and effectively accelerates rendering, especially for highly complex scenes. Furthermore, we demonstrate how level of detail can be integrated into our proposed method to further enhance rendering speed for large-scale scenes. With an optimized implementation, we highlight key practical considerations and thoroughly evaluate the proposed technique and its impact on desktop and mobile devices.
Jonathan Haberl, Philipp Fleck, Clemens Arth
Comput. Graph.3
2025 Assisted Trailer Parking using a Reverse Camera System and Inverse Kinematics
abstract
We introduce a rear-view camera system designed for trailers, comprising multiple cameras, to address the challenges of limited visibility and the awkward maneuverability of a trailer when performing a reverse driving task (e.g., parking). The cameras are placed on the top edge of the trailer looking downwards, while the different camera feeds are transformed and stitched to create a top-down view. The actual steering angle of the vehicle and the hitch angle are measured using IoT hardware. The path of the trailer is determined using inverse kinematics of the vehicle and trailer, and lines are overlaid on the top-down view to provide drivers with dynamic, real-time guiding lines on their smartphone. We describe a practical implementation of our system on one instance of a vehicle and a trailer, and extrapolate the concept to different types of trailers in simulations.
Daniel Kreimer, Philipp Fleck, Thomas Kernbauer, Clemens Arth
IV4
2025 Spatial Augmented Reality for Heavy Machinery Using Laser Projections
Maximilian Tschulik, Thomas Kernbauer, Philipp Fleck, Clemens Arth
Comput. Graph.4
2024 Instant Segmentation and Fitting of Excavations in Subsurface Utility Engineering
abstract
Using augmented reality for subsurface utility engineering (SUE) has benefited from recent advances in sensing hardware, enabling the first practical and commercial applications. However, this progress has uncovered a latent problem - the insufficient quality of existing SUE data in terms of completeness and accuracy. In this work, we present a novel approach to automate the process of aligning existing SUE databases with measurements taken during excavation works, with the potential to correct the deviation from the as-planned to as-built documentation, which is still a big challenge for traditional workers at sight. Our segmentation algorithm performs infrastructure segmentation based on the live capture of an excavation on site. Our fitting approach correlates the inferred position and orientation with the existing digital plan and registers the as-planned model into the as-built state. Our approach is the first to circumvent tedious postprocessing, as it corrects data online and on-site. In our experiments, we show the results of our proposed method on both synthetic data and a set of real excavations.
Marco Stranner, Philipp Fleck, Dieter Schmalstieg, Clemens Arth
IEEE Trans. Vis. Comput. Graph.4
2023 Compact World Anchors: Registration Using Parametric Primitives as Scene Description
abstract
We present a registration method relying on geometric constraints extracted from parametric primitives contained in 3D parametric models. Our method solves the registration in closed-form from three line-to-line, line-to-plane or plane-to-plane correspondences. The approach either works with semantically segmented RGB-D scans of the scene or with the output of plane detection in common frameworks like ARKit and ARCore. Based on the primitives detected in the scene, we build a list of descriptors using the normals and centroids of all the found primitives, and match them against the pre-computed list of descriptors from the model in order to find the scene-to-model primitive correspondences. Finally, we use our closed-form solver to estimate the 6DOFtransformation from three lines and one point, which we obtain from the parametric representations of the model and scene parametric primitives. Quantitative and qualitative experiments on synthetic and real-world data sets demonstrate the performance and robustness of our method. We show that it can be used to create compact world anchors for indoor localization in AR applications on mobile devices leveraging commercial SLAM capabilities.
Fernando Reyes-Aviles, Philipp Fleck, Dieter Schmalstieg, Clemens Arth
IEEE Trans. Vis. Comput. Graph.4
2023 Bag of World Anchors for Instant Large-Scale Localization
abstract
In this work, we present a novel scene description to perform large-scale localization using only geometric constraints. Our work extends compact world anchors with a search data structure to efficiently perform localization and pose estimation of mobile augmented reality devices across multiple platforms (e.g., HoloLens 2, iPad). The algorithm uses a bag-of-words approach to characterize distinct scenes (e.g., rooms). Since the individual scene representations rely on compact geometric (rather than appearance-based) features, the resulting search structure is very lightweight and fast, lending itself to deployment on mobile devices. We present a set of experiments demonstrating the accuracy, performance and scalability of our novel localization method. In addition, we describe several use cases demonstrating how efficient cross-platform localization facilitates sharing of augmented reality experiences.
Fernando Reyes-Aviles, Philipp Fleck, Dieter Schmalstieg, Clemens Arth
IEEE Trans. Vis. Comput. Graph.4
2021 Augmented Reality for Subsurface Utility Engineering, Revisited
abstract
Civil engineering is a primary domain for new augmented reality technologies. In this work, the area of subsurface utility engineering is revisited, and new methods tackling well-known, yet unsolved problems are presented. We describe our solution to the outdoor localization problem, which is deemed one of the most critical issues in outdoor augmented reality, proposing a novel, lightweight hardware platform to generate highly accurate position and orientation estimates in a global context. Furthermore, we present new approaches to drastically improve realism of outdoor data visualizations. First, a novel method to replace physical spray markings by indistinguishable virtual counterparts is described. Second, the visualization of 3D reconstructions of real excavations is presented, fusing seamlessly with the view onto the real environment. We demonstrate the power of these new methods on a set of different outdoor scenarios.
Lasse H. Hansen, Philipp Fleck, Marco Stranner, Dieter Schmalstieg, Clemens Arth
IEEE Trans. Vis. Comput. Graph.5
2019 Towards SLAM-Based Outdoor Localization using Poor GPS and 2.5D Building Models
abstract
In this paper, we address the topic of outdoor localization and tracking using monocular camera setups with poor GPS priors. We leverage 2.5D building maps, which are freely available from open-source databases such as OpenStreetMap. The main contributions of our work are a fast initialization method and a non-linear optimization scheme. The initialization upgrades a visual SLAM reconstruction with an absolute scale. The non-linear optimization uses the 2.5D building model footprint, which further improves the tracking accuracy and the scale estimation. A pose optimization step relates the vision-based camera pose estimation from SLAM to the position information received through GPS, in order to fix the common problem of drift. We evaluate our approach on a set of challenging scenarios. The experimental results show that our approach achieves improved accuracy and robustness with an advantage in run-time over previous setups.
Ruyu Liu, Jianhua Zhang 0002, Shengyong Chen, Clemens Arth
ISMAR4
2018 Measurement Uncertainty Analysis of a Robotic Total Station Simulation
abstract
The design of interactive algorithms for robotic total stations often requires hardware-in-the-Ioop setups during software development and verification. The use of real-time simulation setups can reduce the development and test effort significantly. However, the analysis of the simulation uncertainty is crucial for proper design of simulation setups and for the interpretation of simulation results. In this paper, we present a real-time simulation method for modern robotic total stations. We provide details for an exemplary robotic total station including models of geometry, actuators and sensors. The simulation uncertainty was estimated analytically and verified by Monte Carlo experiments.
Christoph Klug, Clemens Arth, Dieter Schmalstieg, Thomas Gloor
IECON2
2018 Semi-Automatic Registration of a Robotic Total Station and a CAD Model Without Control Points
abstract
The accurate registration of a robotic total station with respect to a given CAD model is a crucial task in the construction industry. Common registration techniques rely on a reference network of control points in the CAD model. One must establish correspondences between control points in the CAD model and measured points in the field. Usually physical markers or natural points of interest are selected as control points. We present a user-guided algorithm for simple and efficient registration of a robotic total station with a CAD model in indoor environments without the need for control points. The user interaction is reduced to selecting a local Manhattan-like corner structure for initial model alignment; accurate registration of the device is carried out automatically. Our algorithm relies on angle and distance measurements only and, therefore, is not limited to vision based robotic total stations. In particular, we propose a new algorithm for robust Manhattan corner extraction.
Christoph Klug, Clemens Arth, Dieter Schmalstieg, Thomas Gloor
IECON2
2018 Efficient Physics-Based Implementation for Realistic Hand-Object Interaction in Virtual Reality
abstract
We propose an efficient physics-based method for dexterous `real hand' - `virtual object' interaction in Virtual Reality environments. Our method is based on the Coulomb friction model, and we show how to efficiently implement it in a commodity VR engine for realtime performance. This model enables very convincing simulations of many types of actions such as pushing, pulling, grasping, or even dexterous manipulations such as spinning objects between fingers without restrictions on the objects' shapes or hand poses. Because it is an analytic model, we do not require any prerecorded data, in contrast to previous methods. For the evaluation of our method, we conduction a pilot study that shows that our method is perceived more realistic and natural, and allows for more diverse interactions. Further, we evaluate the computational complexity of our method to show real-time performance in VR environments.
Markus Höll, Markus Oberweger, Clemens Arth, Vincent Lepetit
VR3
2018 Incremental Structural Modeling Based on Geometric and Statistical Analyses
abstract
Finding high-level semantic information from a point cloud is a challenging task, and it can be used in various applications. For instance, it is useful to compactly represent the scene structure and efficiently understand the scene context. This task is even more challenging when using a hand-held monocular visual SLAM system that outputs a noisy sparse point cloud. In order to tackle this issue, we propose an incremental primitive modeling method using both geometric and statistical analyses for such point cloud. The main idea is to select only reliably-modeled shapes by analyzing the geometric relationship between the point cloud and the estimated shapes. Besides that, a statistical evaluation is incorporated to filter wrongly-detected primitives in a noisy point cloud. As a result of this processing, our approach largely improved precision when compared with state of the art methods. We also show the impact of segmenting and representing a scene using primitives instead of a point cloud.
Rafael Alves Roberto, Joao Paulo Silva do Monte Lima, Hideaki Uchiyama, Clemens Arth, Veronica Teichrieb, Rin-Ichiro Taniguchi, Dieter Schmalstieg
WACV4
2016 Efficient Verification of Holograms Using Mobile Augmented Reality
abstract
Paper documents such as passports, visas and banknotes are frequently checked by inspection of security elements. In particular, optically variable devices such as holograms are important, but difficult to inspect. Augmented Reality can provide all relevant information on standard mobile devices. However, hologram verification on mobiles still takes long and provides lower accuracy than inspection by human individuals using appropriate reference information. We aim to address these drawbacks by automatic matching combined with a special parametrization of an efficient goal-oriented user interface which supports constrained navigation. We first evaluate a series of similarity measures for matching hologram patches to provide a sound basis for automatic decisions. Then a re-parametrized user interface is proposed based on observations of typical user behavior during document capture. These measures help to reduce capture time to approximately 15 s with better decisions regarding the evaluated samples than what can be achieved by untrained users.
Andreas Hartl, Clemens Arth, Jens Grubert, Dieter Schmalstieg
IEEE Trans. Vis. Comput. Graph.2
2015 An Efficient Minimal Solution for Multi-camera Motion
abstract
Summary form only given. We propose an efficient method for estimating the motion of a multi-camera rig from a minimal set of feature correspondences. Existing methods for solving the multi-camera relative pose problem require extra correspondences, are slow to compute, and/or produce a multitude of solutions. Our solution uses a first-order approximation to relative pose in order to simplify the problem and produce an accurate estimate quickly. The solver is applicable to sequential multi-camera motion estimation and is fast enough for real-time implementation in a random sampling framework. Our experiments show that our approach is both stable and efficient on challenging test sequences.
Jonathan Ventura, Clemens Arth, Vincent Lepetit
ICCV2
2015 Tutorial 1: Global-scale Localization in Outdoor Environments for AR
abstract
In this tutorial we aim for a review of existing technologies to perform outdoor localization in urban environments at a global level in full 6DOF using visual sensors primarily. The goal is to provide a clear overview about the current state-of-the-art in global positioning and orientation estimation, which includes a wide range of methods and algorithms from both the Computer Vision and the Augmented Reality community. The main focus is put on methods that are real-time capable, or can at least be applied through a server-client infrastructure. Algorithms that are based on single images, panoramic images, as well as SLAM maps and sparse point cloud reconstructions from SfM will be discussed, together with mobile hardware considerations.The attendees will acquire an overview about the current landscape of technologies employed to facilitate outdoor localization for AR. The tutorial should enable them to get a feeling for the current state-of-the-art of methods for outdoor Augmented Reality.
Clemens Arth, Dieter Schmalstieg
ISMAR1
2015 Tracking and Mapping with a Swarm of Heterogeneous Clients
abstract
In this work, we propose a multi-user system for tracking and mapping, which accommodates mobile clients with different capabilities, mediated by a server capable of providing real-time structure from motion. Clients share their observations of the scene according to their individual capabilities. This can involve only keyframe tracking, but also mapping and map densification, if more computational resources are available. Our contribution is a system architecture that lets heterogeneous clients contribute to a collaborative mapping effort, without prescribing fixed capabilities for the client devices. We investigate the implications that the clients' capabilities have on the collaborative reconstruction effort and its use for AR applications.
Philipp Fleck, Clemens Arth, Christian Pirchheim, Dieter Schmalstieg
ISMAR2
2015 A Particle Filter Approach to Outdoor Localization Using Image-Based Rendering
abstract
We propose an outdoor localization system using a particle filter. In our approach, a textured, geo-registered model of the outdoor environment is used as a reference to estimate the pose of a smartphone. The device position and the orientation obtained from a Global Positioning System (GPS) receiver and an inertial measurement unit (IMU) are used as a first estimation of the true pose. Then, multiple pose hypotheses are randomly distributed about the GPS/IMU measurement and use to produce renderings of the virtual model. With vision-based methods, the rendered images are compared with the image received from the smartphone, and the matching scores are used to update the particle filter. The outcome of our system improves the camera pose estimate in real time without user assistance.
Christian Poglitsch, Clemens Arth, Dieter Schmalstieg, Jonathan Ventura
ISMAR2
2015 Mobile user interfaces for efficient verification of holograms
abstract
Paper documents such as passports, visas and banknotes are frequently checked by inspection of security elements. In particular, view-dependent elements such as holograms are interesting, but the expertise of individuals performing the task varies greatly. Augmented Reality systems can provide all relevant information on standard mobile devices. Hologram verification still takes long and causes considerable load for the user. We aim to address this drawback by first presenting a work flow for recording and automatic matching of hologram patches. Several user interfaces for hologram verification are presented, aiming to noticeably reduce verification time. We evaluate the most promising interfaces in a user study with prototype applications running on off-the-shelf hardware. Our results indicate that there is a significant difference in capture time between interfaces but that users do not prefer the fastest interface.
Andreas Hartl, Jens Grubert, Christian Reinbacher, Clemens Arth, Dieter Schmalstieg
VR4
2015 Instant Outdoor Localization and SLAM Initialization from 2.5D Maps
abstract
We present a method for large-scale geo-localization and global tracking of mobile devices in urban outdoor environments. In contrast to existing methods, we instantaneously initialize and globally register a SLAM map by localizing the first keyframe with respect to widely available untextured 2.5D maps. Given a single image frame and a coarse sensor pose prior, our localization method estimates the absolute camera orientation from straight line segments and the translation by aligning the city map model with a semantic segmentation of the image. We use the resulting 6DOF pose, together with information inferred from the city map model, to reliably initialize and extend a 3D SLAM map in a global coordinate system, applying a model-supported SLAM mapping approach. We show the robustness and accuracy of our localization approach on a challenging dataset, and demonstrate unconstrained global SLAM mapping and tracking of arbitrary camera motion on several sequences.
Clemens Arth, Christian Pirchheim, Jonathan Ventura, Dieter Schmalstieg, Vincent Lepetit
IEEE Trans. Vis. Comput. Graph.1
2014 A Minimal Solution to the Generalized Pose-and-Scale Problem
abstract
We propose a novel solution to the generalized camera pose problem which includes the internal scale of the generalized camera as an unknown parameter. This further generalization of the well-known absolute camera pose problem has applications in multi-frame loop closure. While a well-calibrated camera rig has a fixed and known scale, camera trajectories produced by monocular motion estimation necessarily lack a scale estimate. Thus, when performing loop closure in monocular visual odometry, or registering separate structure-from-motion reconstructions, we must estimate a seven degree-of-freedom similarity transform from corresponding observations. Existing approaches solve this problem, in specialized configurations, by aligning 3D triangulated points or individual camera pose estimates. Our approach handles general configurations of rays and points and directly estimates the full similarity transformation from the 2D-3D correspondences. Four correspondences are needed in the minimal case, which has eight possible solutions. The minimal solver can be used in a hypothesize-and-test architecture for robust transformation estimation. Our solver also produces a least-squares estimate in the overdetermined case. The approach is evaluated experimentally on synthetic and real datasets, and is shown to produce higher accuracy solutions to multi-frame loop closure than existing approaches.
Jonathan Ventura, Clemens Arth, Gerhard Reitmayr, Dieter Schmalstieg
CVPR2
2014 Global Localization from Monocular SLAM on a Mobile Phone
abstract
We propose the combination of a keyframe-based monocular SLAM system and a global localization method. The SLAM system runs locally on a camera-equipped mobile client and provides continuous, relative 6DoF pose estimation as well as keyframe images with computed camera locations. As the local map expands, a server process localizes the keyframes with a pre-made, globally-registered map and returns the global registration correction to the mobile client. The localization result is updated each time a keyframe is added, and observations of global anchor points are added to the client-side bundle adjustment process to further refine the SLAM map registration and limit drift. The end result is a 6DoF tracking and mapping system which provides globally registered tracking in real-time on a mobile device, overcomes the difficulties of localization with a narrow field-of-view mobile phone camera, and is not limited to tracking only in areas covered by the offline reconstruction.
Jonathan Ventura, Clemens Arth, Gerhard Reitmayr, Dieter Schmalstieg
IEEE Trans. Vis. Comput. Graph.2
2013 Panoramic mapping on a mobile phone GPU
abstract
Creating panoramic images in real-time is an expensive operation for mobile devices. Mapping of individual pixels into the panoramic image is the main focus of this paper, since it is one of the most time consuming parts. The pixel-mapping process is transferred from the Central Processing Unit (CPU) to the Graphics Processing Unit (GPU). The independence of pixels being projected allows OpenGL shaders to perform this operation very efficiently. We propose a shader-based mapping approach and confront it with an existing solution. The application is implemented for Android phones and works fluently on current generation devices.
Georg Reinisch, Clemens Arth, Dieter Schmalstieg
ISMAR2
2012 Full 6DOF Pose Estimation from Geo-Located Images
Clemens Arth, Gerhard Reitmayr, Dieter Schmalstieg
ACCV (3)1
2012 Exploiting sensors on mobile phones to improve wide-area localization
Clemens Arth, Alessandro Mulloni, Dieter Schmalstieg
ICPR1
2011 Real-time self-localization from panoramic images on mobile devices
abstract
Self-localization in large environments is a vital task for accurately registered information visualization in outdoor Augmented Reality (AR) applications. In this work, we present a system for self-localization on mobile phones using a GPS prior and an online-generated panoramic view of the user's environment. The approach is suitable for executing entirely on current generation mobile devices, such as smartphones. Parallel execution of online incremental panorama generation and accurate 6DOF pose estimation using 3D point reconstructions allows for real-time self-localization and registration in large-scale environments. The power of our approach is demonstrated in several experimental evaluations.
Clemens Arth, Manfred Klopschitz, Gerhard Reitmayr, Dieter Schmalstieg
ISMAR1
2011 Rapid scene reconstruction on mobile phones from panoramic images
abstract
Rapid 3D reconstruction of environments has become an active research topic due to the importance of 3D models in a huge number of applications, be it in Augmented Reality (AR), architecture or other commercial areas. In this paper we present a novel system that allows for the generation of a coarse 3D model of the environment within several seconds on mobile smartphones. By using a very fast and flexible algorithm a set of panoramic images is captured to form the basis of wide field-of-view images required for reliable and robust reconstruction. A cheap on-line space carving approach based on Delaunay triangulation is employed to obtain dense, polygonal, textured representations. The use of an intuitive method to capture these images, as well as the efficiency of the reconstruction approach allows for an application on recent mobile phone hardware, giving visually pleasing results almost instantly.
Qi Pan, Clemens Arth, Edward Rosten, Gerhard Reitmayr, Tom Drummond
ISMAR2
2009 Wide area localization on mobile phones
abstract
We present a fast and memory efficient method for localizing a mobile user's 6DOF pose from a single camera image. Our approach registers a view with respect to a sparse 3D point reconstruction. The 3D point dataset is partitioned into pieces based on visibility constraints and occlusion culling, making it scalable and efficient to handle. Starting with a coarse guess, our system only considers features that can be seen from the user's position. Our method is resource efficient, usually requiring only a few megabytes of memory, thereby making it feasible to run on low-end devices such as mobile phones. At the same time it is fast enough to give instant results on this device class.
Clemens Arth, Daniel Wagner 0003, Manfred Klopschitz, Arnold Irschara, Dieter Schmalstieg
ISMAR1
2007 Detecting, Tracking and Recognizing License Plates
Michael Donoser, Clemens Arth, Horst Bischof
ACCV (2)2
2007 Real-Time License Plate Recognition on an Embedded DSP-Platform
abstract
In this paper we present a full-featured license plate detection and recognition system. The system is implemented on an embedded DSP platform and processes a video stream in real-time. It consists of a detection and a character recognition module. The detector is based on the AdaBoost approach presented by Viola and Jones. Detected license plates are segmented into individual characters by using a region-based approach. Character classification is performed with support vector classification. In order to speed up the detection process on the embedded device, a Kalman tracker is integrated into the system. The search area of the detector is limited to locations where the next location of a license plate is predicted. Furthermore, classification results of subsequent frames are combined to improve the class accuracy. The major advantages of our system are its real-time capability and that it does not require any additional sensor input (e.g. from infrared sensors) except a video stream. We evaluate our system on a large number of vehicles and license plates using bad quality video and show that the low resolution can be partly compensated by combining classification results of subsequent frames.
Clemens Arth, Florian Limberger, Horst Bischof
CVPR1
2007 Robust Local Features and their Application in Self-Calibration and Object Recognition on Embedded Systems
abstract
In recent years many powerful computer vision algorithms have been invented, making automatic or semiautomatic solutions to many popular vision tasks, such as visual object recognition or camera calibration, possible. On the other hand embedded vision platforms and solutions such as smart cameras have successfully emerged, however, only offering limited computational and memory resources. The first contribution of this paper is the investigation of a set of robust local feature detectors and descriptors for application on embedded systems. We briefly describe the methods involved, i.e. the DoG (difference of Gaussian) and MSER (maximally stable extremal regions) detector as well as the PCA-SIFT descriptor, and discuss their suitability for smart systems and their qualification for given tasks. The second contribution of this work is the experimental evaluation of these methods on two challenging tasks, namely fully embedded object recognition on a moderate size database and on the task of robust camera calibration. Our approach is fortified by encouraging results we present at length.
Clemens Arth, Christian Leistner, Horst Bischof
CVPR1
2007 Dual-Layer Visual Vocabulary Tree Hypotheses for Object Recognition
abstract
This paper introduces an efficient method to substantially increase the recognition performance of a vocabulary tree based recognition system. We propose to enhance the hypothesis obtained by a standard inverse object voting algorithm with reliable descriptor co-occurrences. The algorithm operates on different layers of a standard k-means tree benefiting from the advantages of different levels of information abstraction. The visual vocabulary tree shows good results when a large number of distinctive descriptors form a large visual vocabulary. Co-occurrences perform well even on a coarse object representation with a small number of visual words. An arbitration strategy with minimal computational effort combines the specific strengths of the particular representations. We demonstrate the achieved performance boost and robustness to occlusions in a challenging object recognition task.
Sandra Ober, Clemens Arth, Horst Bischof
ICIP (6)3