VLDB 2026 Research / reviewers in the wild / expert
Margarita Chli
dblp:69/2994
· DBLP profile ↗
51ranked-venue papers
2as first author
15since 2021 · last 2024
0000-0001-5611-7492ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 2 first-author · 15 since 2021Systems, architecture and hardware · 38 · 1 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Hyperion - A Fast, Versatile Symbolic Gaussian Belief Propagation Framework for Continuous-Time SLAM
David Hug, Ignacio Alzugaray, Margarita Chli |
ECCV (30) | 3 |
| 2024 | Aerial Image-based Inter-day Registration for Precision AgricultureabstractSatellite imagery has traditionally been used to collect crop statistics, but its low resolution and registration accuracy limit agricultural analytics to plant stand levels and large areas. Precision agriculture seeks analytic tools at near single plant level, and this work explores how to improve aerial photogrammetry to enable inter-day precision agriculture analytics for intervals of up to a month.Our work starts by presenting an accurately registered image time series, captured up to twice a week, by an unmanned aerial vehicle over a wheat crop field. The dataset is registered using photogrammetry aided by fiducial ground control points (GCPs). Unfortunately, GCPs severely disrupt crop management activities. To address this, we propose a novel inter-day registration approach that only relies once on GCPs, at the beginning of the season.The method utilises LoFTR [1], a state-of-the-art image-matching transformer. The original LoFTR network was trained using imagery of outdoor urban areas. One of our contributions is to extend LoFTR’s training method, which uses matching images of a static scene, to a dynamic scene of plants undergoing growth. Another contribution is a thorough evaluation of our registration method that integrates intraday crop reconstruction with earlier-day scans in a seven degree-of-freedom alignment. Experimental results show the advantage of our approach over other matching algorithms and demonstrate the importance of retraining using crop scenes, and a training method customised for growing crops, with an average registration error of 27 cm across a season. Franz Daxinger, Lukas Roth, Fabiola Maffra, Paul Beardsley, Margarita Chli, Lucas Teixeira |
ICRA | 6 |
| 2024 | Temporal- and Viewpoint-Invariant Registration for Under-Canopy Footage using Deep-Learning-based Bird's-Eye View PredictionabstractConducting visual assessments under the canopy using mobile robots is an emerging task in smart farming and forestry. However, it is challenging to register images across different data-collection days, especially across seasons, due to the self-occluding geometry and temporal dynamics in forests and orchards. This paper proposes a new approach for registering under-canopy image sequences in general and in these situations. Our methodology leverages standard GPS data and deep-learning-based perspective to bird’s-eye view conversion to provide an initial estimation of the positions of the trees in images and their association across datasets. Furthermore, it introduces an innovative strategy for extracting tree trunks and clean ground surfaces from noisy and sparse 3D reconstructions created from the image sequences, utilizing these features to achieve precise alignment. Our robust alignment method effectively mitigates position and scale drift, which may arise from GPS inaccuracies and Sparse Structure from Motion (SfM) limitations. We evaluate our approach on three challenging real-world datasets, demonstrating that our method outperforms ICP-based methods on average by 50%, and surpasses FGR and TEASER++ by over 90% in alignment accuracy. These results highlight our method’s cost efficiency and robustness, even in the presence of severe outliers and sparsity. https://github.com/VIS4ROB-lab/bev_undercanopy_registration Jiawei Zhou 0004, Ruben Mascaro, Cesar Dario Cadena Lerma, Margarita Chli, Lucas Teixeira |
IROS | 4 |
| 2024 | Real-Time Semantic Segmentation in Natural Environments with SAM-assisted Sim-to-Real Domain TransferabstractSemantic segmentation plays a pivotal role in many robotic applications requiring high-level scene understanding, such as smart farming, where the precise identification of trees or plants can aid navigation and crop monitoring tasks. While deep-learning-based semantic segmentation approaches have reached outstanding performance in recent years, they demand large amounts of labeled data for training. Inspired by modern Unsupervised Domain Adaptation (UDA) techniques, in this paper, we introduce a two-step training pipeline specifically tailored to challenging natural scenes, where the availability of annotated data is often quite limited. Our strategy involves the initial training of a powerful domain adaptive architecture, followed by a refinement stage, where segmentation masks predicted by the Segment Anything Model (SAM) are used to improve the accuracy of the predictions on the target dataset. These refined predictions serve as pseudo-labels to supervise the training of a final distilled architecture for real-time deployment. Extensive experiments conducted in two real-world scenes demonstrate the effectiveness of the proposed method. Specifically, we show that our pipeline enables the training of a MobileNetV3 that achieves significant mIoU gains of 3.60% and 11.40% on our two datasets compared to the DAFormer while only demanding 1/15 of the latter’s inference time. Code and datasets are available at https://github.com/VIS4ROB-lab/nature_uda_rt_segmentation. Ruben Mascaro, Margarita Chli, Lucas Teixeira |
IROS | 3 |
| 2023 | Domain-Adaptive Semantic Segmentation with Memory-Efficient Cross-Domain Transformers
Ruben Mascaro, Lucas Teixeira, Margarita Chli |
BMVC | 3 |
| 2023 | Cross-Agent Relocalization for Decentralized Collaborative SLAMabstractState-of-the-art decentralized collaborative Simultaneous Localization And Mapping (SLAM) systems crucially lack the ability to effectively use well-mapped areas generated by other agents in the team for relocalization. This often leads to map redundancy between agents, inefficient communication, and the need for costly re-mapping of areas previously mapped by other agents. In this work, we propose a strategy to efficiently share the areas mapped by different agents in a collaborative, decentralized SLAM system. This approach directly addresses map redundancy while maintaining the consistency of the estimates across the agents and keeping the overall system scalable in terms of cross-agent communication and individual computational effort. Our method leverages covisibility information between keyframes instantiated by different agents to transfer local sub-maps on-the-fly in a completely decentralized, peer-to-peer fashion. A globally consistent estimate is achieved by solving a distributed bundle adjustment problem using the Alternating Direction Method of Multipliers (ADMM), where we enforce constraints on shared map points and keyframes across agents. Philipp Bänninger, Ignacio Alzugaray, Marco Karrer, Margarita Chli |
ICRA | 4 |
| 2023 | Target-Aware Implicit Mapping for Agricultural Crop InspectionabstractCrop inspection is a critical part of modern agricultural practices that helps farmers assess the current status of a field and then make crop management decisions. Current crop inspection methods are labour-intensive tasks, which makes them rather slow and expensive to apply. In this paper, we exploit recent advancements in implicit mapping to tackle the challenging context of agricultural environments to create dense maps of crop rows with high enough fidelity to be useful for automated crop inspection. Specifically, we map strawberry and sweet pepper crop rows using RGB images captured by a wheeled mobile field robot inside a greenhouse and then use this data to build 3D maps to document the development of plants and fruits. Our Target-Aware Implicit Mapping system (TAIM) uses a SLAM-based pose initialization strategy for robust pose convergence, an efficient information-guided training sample selection framework for faster loss reduction, and focuses on exploiting training samples for fruit regions of the scene, which are critical for crop inspection tasks, to create more accurate maps in less time. Shane Kelly, Alessandro Riccardi, Elias Marks, Federico Magistri, Tiziano Guadagnino, Margarita Chli, Cyrill Stachniss |
ICRA | 6 |
| 2023 | COVINS-G: A Generic Back-end for Collaborative Visual-Inertial SLAMabstractCollaborative SLAM is at the core of perception in multi-robot systems as it enables the co-localization of the team of robots in a common reference frame, which is of vital importance for any coordination amongst them. The paradigm of a centralized architecture is well established, with the robots (i.e. agents) running Visual-Inertial Odometry (VIO) onboard while communicating relevant data, such as e.g. Keyframes (KFs), to a central back-end (i.e. server), which then merges and optimizes the joint maps of the agents. While these frameworks have proven to be successful, their capability and performance are highly dependent on the choice of the VIO front-end, thus limiting their flexibility. In this work, we present COVINSG, a generalized back-end building upon the COVINS [1] framework, enabling the compatibility of the server-back-end with any arbitrary VIO front-end, including, for example, off-the-shelf cameras with odometry capabilities, such as the Realsense T265. The COVINS-G back-end deploys a multi-camera relative pose estimation algorithm for computing the loop-closure constraints allowing the system to work purely on 2D image data. In the experimental evaluation, we show on-par accuracy with state-of-the-art multi-session and collaborative SLAM systems, while demonstrating the flexibility and generality of our approach by employing different front-ends onboard collaborating agents within the same mission. The COVINS-G codebase along with a generalized front-end wrapper to allow any existing VIO front-end to be readily used in combination with the proposed collaborative back-end is open-sourced. Video- https://youtu.be/FoJfXCfaYDw Manthan Patel, Marco Karrer, Philipp Bänninger, Margarita Chli |
ICRA | 4 |
| 2023 | Decentralised Multi-Robot Exploration Using Monte Carlo Tree SearchabstractAutonomous robotic systems are useful in automating tasks such as inspection and surveying of unknown areas, where speed is often an important factor. In order to effectively reduce the time required to complete missions, an efficient exploration and coordination strategy is needed. In this spirit, this work proposes an approach based on the Monte Carlo Tree Search (MCTS) algorithm to guide robots during exploration missions. Our method first expands a search tree of possible actions from the robot's position towards unknown regions, and then selects the sequence of movements that best drive the exploration process forward with respect to a given reward function. The proposed approach, which is able to balance short- and long-term decision-making, is then extended to accommodate the presence of multiple robots, in a bid to push the efficiency of exploration further. Our method allows for the coordination of the robots' movements in a decentralized manner, relying on point-to-point communication. This results in an efficient strategy, which we refer to as Decentralized Monte Carlo Exploration (DMCE). The experimental results demonstrate that our pipeline outperforms a greedy exploration approach, as well as state-of-the-art planners, with up to 30% reduction in exploration times in a series of real-world maps. Sean Bone, Luca Bartolomei 0002, Florian Kennel-Maushart, Margarita Chli |
IROS | 4 |
| 2022 | Autonomous Emergency Landing for Multicopters using Deep Reinforcement LearningabstractThis work presents a pipeline for autonomous emergency landing for multicopters, such as rotary wing Unmanned Aerial Vehicles (UAVs), using deep Reinforcement Learning (RL). Mechanical malfunctions, strong winds, sudden battery life drops (e.g, due to cold weather), failure in localization or GPS jamming are not uncommon and all constitute emergency situations that require a UAV to abort its mission early and land as quickly as possible in the immediate vicinity. To this end, it is crucial for a UAV that is deployed in real missions to be able to detect a safe landing spot efficiently and proceed to land autonomously, avoiding damage to both its integrity and the surroundings. Driven by the advances in semantic segmentation and depth completion using machine learning, the proposed architecture uses deep RL to infer actions from semantic and depth information, flying the robot towards secure areas, while respecting safety constraints. Thanks to our robust training strategy and the choice of these mid-level representations as input to the RL agent, we show that our policy can directly transfer to the real world, without the need for any additional fine-tuning. In a series of challenging experiments both in simulation and with a real platform, we demonstrate that our planner guides a rotorcraft UAV to a safe landing spot up to 1.5 times faster and with double success rate than the state of the art (including a commercially available solution), paving the way towards realistically deployable UAVs. Luca Bartolomei 0002, Yves Kompis, Lucas Teixeira, Margarita Chli |
IROS | 4 |
| 2022 | T-PRM: Temporal Probabilistic Roadmap for Path Planning in Dynamic EnvironmentsabstractSampling-based motion planners are widely used in robotics due to their simplicity, flexibility and computational efficiency. However, in their most basic form, these algorithms operate under the assumption of static scenes and lack the ability to avoid collisions with dynamic (i.e. moving) obstacles. This raises safety concerns, limiting the range of possible applications of mobile robots in the real world. Motivated by these challenges, in this work we present Temporal-PRM, a novel sampling-based path-planning algorithm that performs obstacle avoidance in dynamic environments. The proposed approach extends the original Probabilistic Roadmap (PRM) with the notion of time, generating an augmented graph-like structure that can be efficiently queried using a time-aware variant of the A* search algorithm, also introduced in this paper. Our design maintains all the properties of PRM, such as the ability to perform multiple queries and to find smooth paths, while circumventing its downside by enabling collision avoidance in highly dynamic scenes with a minor increase in the computational cost. Through a series of challenging experiments in highly cluttered and dynamic environments, we demonstrate that the proposed path planner outperforms other state-of-the-art sampling-based solvers. Moreover, we show that our algorithm can run onboard a flying robot, performing obstacle avoidance in real time. Matthias Hüppi, Luca Bartolomei 0002, Ruben Mascaro, Margarita Chli |
IROS | 4 |
| 2022 | Voxfield: Non-Projective Signed Distance Fields for Online Planning and 3D ReconstructionabstractCreating accurate maps of complex, unknown environments is of utmost importance for truly autonomous navigation robot. However, building these maps online is far from trivial, especially when dealing with large amounts of raw sensor readings on a computation and energy constrained mobile system, such as a small drone. While numerous approaches tackling this problem have emerged in recent years, the mapping accuracy is often sacrificed as systematic approximation errors are tolerated for efficiency's sake. Motivated by these challenges, we propose Voxfield, a mapping framework that can generate maps online with higher accuracy and lower computational burden than the state of the art. Built upon the novel formulation of non-projective truncated signed distance fields (TSDFs), our approach produces more accurate and complete maps, suitable for surface reconstruction. Additionally, it enables efficient generation of Euclidean signed distance fields (ESDFs), useful e.g., for path planning, that does not suffer from typical approximation errors. Through a series of experiments with public datasets, both real-world and synthetic, we demonstrate that our method beats the state of the art in map coverage, accuracy and computational time. Moreover, we show that Voxfield can be utilized as a back-end in recent multi-resolution mapping frameworks, producing high quality maps even in large-scale experiments. Finally, we validate our method by running it onboard a quadrotor, showing it can generate accurate ESDF maps usable for real-time path planning and obstacle avoidance. Yue Pan 0009, Yves Kompis, Luca Bartolomei 0002, Ruben Mascaro, Cyrill Stachniss, Margarita Chli |
IROS | 6 |
| 2021 | Distributed Variable-Baseline Stereo SLAM from two UAVsabstractVisual-Inertial Odometry (VIO) has been widely used and researched to control and aid the automation of navigation of robots especially in the absence of absolute position measurements, such as GPS. However, when the observable landmarks in the scene lie far away, as in high-altitude flights for example, the fidelity of the metric scale estimate in VIO greatly degrades. Aiming to tackle this issue, in this work, we utilize the virtual stereo setup formed by two Unmanned Aerial Vehicles (UAVs), equipped with one camera and one Inertial Measurement Unit (IMU) each, exploiting their view overlap and relative distance measurements between them using onboard Ultra-Wideband (UWB) modules to enable collaborative VIO. In particular, we propose a decentralized collaborative estimation scheme, where each agent holds its own local map, achieving a low pose estimation latency, while ensuring consistency of each agents’ estimates via consensus-based optimization. Following a thorough evaluation in photorealistic simulations, we demonstrate the effectiveness of the approach at high-altitude flights of up to 160m, going significantly beyond the capabilities of state-of-the-art VIO methods. Finally, we show the advantage of actively adjusting the baseline on-the-fly over a fixed, target baseline, resulting in a significant reduction of the estimation error. Marco Karrer, Margarita Chli |
ICRA | 2 |
| 2021 | Diffuser: Multi-View 2D-to-3D Label Diffusion for Semantic Scene SegmentationabstractSemantic 3D scene understanding is a fundamental problem in computer vision and robotics. Despite recent advances in deep learning, its application to multi-domain 3D semantic segmentation typically suffers from the lack of extensive enough annotated 3D datasets. On the contrary, 2D neural networks benefit from existing large amounts of training data and can be applied to a wider variety of environments, sometimes even without need for retraining. In this paper, we present ‘Diffuser’, a novel and efficient multi-view fusion framework that leverages 2D semantic segmentation of multiple image views of a scene to produce a consistent and refined 3D segmentation. We formulate the 3D segmentation task as a transductive label diffusion problem on a graph, where multi-view and 3D geometric properties are used to propagate semantic labels from the 2D image space to the 3D map. Experiments conducted on indoor and outdoor challenging datasets demonstrate the versatility of our approach, as well as its effectiveness for both global 3D scene labeling and single RGB-D frame segmentation. Furthermore, we show a significant increase in 3D segmentation accuracy compared to probabilistic fusion methods employed in several state-of-the-art multi-view approaches, with little computational overhead. Ruben Mascaro, Lucas Teixeira, Margarita Chli |
ICRA | 3 |
| 2021 | Semantic-aware Active Perception for UAVs using Deep Reinforcement LearningabstractThis work presents a semantic-aware path-planning pipeline for Unmanned Aerial Vehicles (UAVs) using deep reinforcement learning for vision-based navigation in challenging environments. Driven by the maturity of works in semantic segmentation, the proposed path-planning architecture uses reinforcement learning to distinguish the parts of the scene that are perceptually more informative using semantic cues, in effect guiding more robust, repeatable, and accurate navigation of the UAV to the predefined goal destination. Assuming that the UAV performs vision-based state estimation, such as keyframe-based visual odometry, and semantic segmentation onboard, the proposed deep policy network continuously evaluates the optimal relative perceptual informativeness of each semantic class in view. A perception-aware path planner uses these informativeness values to perform trajectory optimization in order to generate the next best action with respect to the current state and the perception quality of the surroundings, essentially guiding the UAV to avoid flying over perceptually degraded regions. Thanks to the use of semantic cues, the policy can be trained in a large number of non-photorealistic randomly-generated scenes, and results to an architecture that is generalizable to environments with the same semantic classes, independently of their visual appearance. Extensive evaluations on challenging, photorealistic simulations reveal a remarkable improvement in robustness and success rate with the proposed approach over the state of the art in active perception. Video – https://youtu.be/RaO3whUBVnc Luca Bartolomei 0002, Lucas Teixeira, Margarita Chli |
IROS | 3 |
| 2020 | HyperSLAM: A Generic and Modular Approach to Sensor Fusion and Simultaneous Localization And Mapping in Continuous-TimeabstractWithin recent years, Continuous-Time Simultaneous Localization And Mapping (CTSLAM) formalisms have become subject to increased attention from the scientific community due to their vast potential in facilitating motioncorrected feature reprojection and direct unsynchronized multi-rate sensor fusion. They also hold the promise of yielding better estimates in traditional sensor setups (e.g. visual, inertial) when compared to conventional discretetime approaches. Related works mostly rely on cubic, C2-continuous, uniform cumulative B-Splines to exemplify and demonstrate the benefits inherent to continuous-time representations. However, as this type of splines gives rise to continuous trajectories by blending uniformly distributed SE3transformations in time, it is prone to underor overparametrize underlying motions with varying volatility and prohibits dynamic trajectory refinement or sparsification by design. In light of this, we propose employing a more generalized and efficient non-uniform split interpolation method in R×SU2×R3and commence with development of `HyperSLAM', a generic and modular CTSLAM framework. The efficacy of our approach is exemplified in proof-ofconcept simulations based on a visual, monocular setup. David Hug, Margarita Chli |
3DV | 2 |
| 2020 | HASTE: multi-Hypothesis Asynchronous Speeded-up Tracking of Events
Ignacio Alzugaray, Margarita Chli |
BMVC | 2 |
| 2020 | Multi-robot Coordination with Agent-Server Architecture for Autonomous Navigation in Partially Unknown EnvironmentsabstractIn this work, we present a system architecture to enable autonomous navigation of multiple agents across user-selected global interest points in a partially unknown environment. The system is composed of a server and a team of agents, here small aircrafts. Leveraging this architecture, computation-ally demanding tasks, such as global dense mapping and global path planning can be outsourced to a potentially powerful central server, limiting the onboard computation for each agent to local pose estimation using Visual-Inertial Odometry (VIO) and local path planning for obstacle avoidance. By assigning priorities to the agents, we propose a hierarchical multi-robot global planning pipeline, which avoids collisions amongst the agents and computes their paths towards the respective goals. The resulting global paths are communicated to the agents and serve as reference input to the local planner running onboard each agent. In contrast to previous works, here we relax the common assumption of a previously mapped environment and perfect knowledge about the state, and we show the effectiveness of the proposed approach in photo-realistic simulations with up to four agents operating in an industrial environment. Luca Bartolomei 0002, Marco Karrer, Margarita Chli |
IROS | 3 |
| 2020 | Perception-aware Path Planning for UAVs using Semantic SegmentationabstractIn this work, we present a perception-aware path-planning pipeline for Unmanned Aerial Vehicles (UAVs) for navigation in challenging environments. The objective is to reach a given destination safely and accurately by relying on monocular camera-based state estimators, such as Keyframe-based Visual-Inertial Odometry (VIO) systems. Motivated by the recent advances in semantic segmentation using deep learning, our path-planning architecture takes into consideration the semantic classes of parts of the scene that are perceptually more informative than others. This work proposes a planning strategy capable of avoiding both texture-less regions and problematic areas, such as lakes and oceans, that may cause large drift or failures in the robot's pose estimation, by using the semantic information to compute the next best action with respect to perception quality. We design a hierarchical planner, composed of an A* path-search step followed by B-Spline trajectory optimization. While the A* steers the UAV towards informative areas, the optimizer keeps the most promising landmarks in the camera's field of view. We extensively evaluate our approach in a set of photo-realistic simulations, showing a remarkable improvement with respect to the state-of-the-art in active perception. Luca Bartolomei 0002, Lucas Teixeira, Margarita Chli |
IROS | 3 |
| 2019 | Asynchronous Multi-Hypothesis Tracking of Features with Event CamerasabstractWith the emergence of event cameras, increasing research effort has been focusing on processing the asynchronous stream of events. With each event encoding a discrete intensity change at a particular pixel, uniquely time-stamped with high accuracy, this sensing information is so fundamentally different to the data provided by traditional frame-based cameras that most of the well-established vision algorithms are not applicable. Inspired by the need of effective event-based tracking, this paper addresses the tracking of generic patch features relying solely on events, while exploiting their asynchronicity and high-temporal resolution. The proposed approach outperforms the state-of-the-art in event-based feature tracking on well-established event camera datasets, retrieving longer and more accurate feature tracks at higher a frequency. Considering tracking as an optimization problem of matching the current view to a feature template, the proposed method implements a simple and efficient technique that only requires the evaluation of a discrete set of tracking hypotheses. Ignacio Alzugaray, Margarita Chli |
3DV | 2 |
| 2019 | On the Redundancy Detection in Keyframe-Based SLAMabstractEgomotion and scene estimation is a key component in automating robot navigation, as well as in virtual reality applications for mobile phones or head-mounted displays. It is well known, however, that with long exploratory trajectories and multi-session mapping for long-term autonomy or collaborative applications, the maintenance of the ever-increasing size of these maps quickly becomes a bottleneck. With the explosion of data resulting in increasing runtime of the optimization algorithms ensuring the accuracy of the Simultaneous Localization And Mapping (SLAM) estimates, the large quantity of collected experiences is imposing hard limits on the scalability of such techniques. Considering the keyframe-based paradigm of SLAM techniques, this paper investigates the redundancy inherent in SLAM maps, by quantifying the information of different experiences of the scene as encoded in keyframes. Here we propose and evaluate different information-theoretic and heuristic metrics to remove dispensable scene measurements with minimal impact on the accuracy of the SLAM estimates. Evaluating the proposed metrics in two state-of-the-art centralized collaborative SLAM systems, we provide our key insights into how to identify redundancy in keyframe-based SLAM. Patrik Schmuck, Margarita Chli |
3DV | 2 |
| 2019 | A Fully-Integrated Sensing and Control System for High-Accuracy Mobile Robotic Building ConstructionabstractWe present a fully-integrated sensing and control system which enables mobile manipulator robots to execute building tasks with millimeter-scale accuracy on building construction sites. The approach leverages multi-modal sensing capabilities for state estimation, tight integration with digital building models, and integrated trajectory planning and whole-body motion control. A novel method for high-accuracy localization updates relative to the known building structure is proposed. The approach is implemented on a real platform and tested under realistic construction conditions. We show that the system can achieve sub-cm end-effector positioning accuracy during fully autonomous operation using solely onboard sensing. Abel Gawel, Roland Siegwart, Marco Hutter 0001, Timothy Sandy, Hermann Blum, Johannes Pankert, Koen Krämer, Luca Bartolomei 0002, Selen Ercan Jenny, Farbod Farshidian, Margarita Chli, Fabio Gramazio |
IROS | 11 |
| 2018 | ACE: An Efficient Asynchronous Corner Tracker for Event CamerasabstractThe emergence of bio-inspired event cameras has opened up new exciting possibilities in high-frequency tracking, overcoming some of the limitations of traditional frame-based vision (e.g. motion blur during high-speed motions or saturation in scenes with high dynamic range). As a result, research has been focusing on the processing of their unusual output: an asynchronous stream of events. With the majority of existing techniques discretizing the event-stream into frame-like representations, we are yet to harness the true power of these cameras. In this paper, we propose the ACE tracker: a purely asynchronous framework to track corner-event features. Evaluation on benchmarking datasets reveals significant improvements in accuracy and computational efficiency in comparison to state-of-the-art event-based trackers. ACE achieves robust performance even in challenging scenarios, where traditional frame-based vision algorithms fail. Ignacio Alzugaray, Margarita Chli |
3DV | 2 |
| 2018 | Learning Deep Descriptors With Scale-Aware Triplet NetworksabstractResearch on learning suitable feature descriptors for Computer Vision has recently shifted to deep learning where the biggest challenge lies with the formulation of appropriate loss functions, especially since the descriptors to be learned are not known at training time. While approaches such as Siamese and triplet losses have been applied with success, it is still not well understood what makes a good loss function. In this spirit, this work demonstrates that many commonly used losses suffer from a range of problems. Based on this analysis, we introduce mixed-context losses and scale-aware sampling, two methods that when combined enable networks to learn consistently scaled descriptors for the first time. Michel Keller, Zetao Chen, Fabiola Maffra, Patrik Schmuck, Margarita Chli |
CVPR | 5 |
| 2018 | Collaborative 6DoF Relative Pose Estimation for Two UAVs with Overlapping Fields of ViewabstractDriven by the promise of leveraging the benefits of collaborative robot operation, this paper presents an approach to estimate the relative transformation between two small Unmanned Aerial Vehicles (UAVs), each equipped with a single camera and an inertial sensor, comprising the first step of any meaningful collaboration. Formation flying and collaborative object manipulation are some of the few tasks that the proposed work has direct applications on, while forming a variable-baseline stereo rig using two UAVs carrying a monocular camera each promises unprecedented effectiveness in collaborative scene estimation. Assuming an overlap in the UAVs' fields of view, in the proposed framework, each UAV runs monocular-inertial odometry onboard, while an Extended Kalman Filter fuses the UAVs' estimates and common image measurements to estimate the metrically scaled relative transformation between them, in realtime. Decoupling the direction of the baseline between the cameras of the two UAVs from its magnitude, this work enables consistent and robust estimation of the uncertainty of the relative pose estimation. Our evaluation on both on simulated data and benchmarking datasets consisting of real aerial data, reveals the power of the proposed methodology in a variety of scenarios. Video - https://youtu.be/AmkkaXa2601. Marco Karrer, Mina Kamel 0001, Roland Siegwart, Margarita Chli |
ICRA | 5 |
| 2018 | Towards Globally Consistent Visual-Inertial Collaborative SLAMabstractMotivated by the need for globally consistent tracking and mapping before autonomous robot navigation becomes realistically feasible, this paper presents a novel backend to monocular-inertial odometry. As some of the most challenging platforms for vision-based perception, we evaluate the performance of our system using Unmanned Aerial Vehicles (UAV s). Our experimental validation demonstrates that the proposed approach achieves drift correction and metric scale estimation from a single UAV on benchmarking datasets. Furthermore, the generality of our approach is demonstrated to achieve globally consistent maps built in a collaborative manner from two UAVs, each equipped with a monocular-inertial sensor suite, showing the possible gains opened by collaboration amongst robots to perform SLAM. Video - https://youtu.be/wbX36HBu2Eg. Marco Karrer, Margarita Chli |
ICRA | 2 |
| 2018 | Viewpoint-Tolerant Place Recognition Combining 2D and 3D Information for UAV NavigationabstractThe booming interest in Unmanned Aerial Vehicles (UAV s) is fed by their potentially great impact, however progress is hindered by their limited perception capabilities. While vision-based odometry was shown to run successfully onboard UAV s, loop-closure detection to correct for drift or to recover from tracking failures, has so far, proven particularly challenging for UAVs. At the heart of this is the problem of viewpoint-tolerant place recognition; in stark difference to ground robots, UAVs can revisit a scene from very different viewpoints. As a result, existing approaches struggle greatly as the task at hand violates underlying assumptions in assessing scene similarity. In this paper, we propose a place recognition framework, which exploits both efficient binary features and noisy estimates of the local 3D geometry, which are anyway computed for visual-inertial odometry onboard the UAV. Attaching both an appearance and a geometry signature to each `location', the proposed approach demonstrates unprecedented recall for perfect precision as well as high quality loop-closing transformations on both flying and hand-held datasets exhibiting large viewpoint and appearance changes as well as perceptual aliasing. Fabiola Maffra, Zetao Chen, Margarita Chli |
ICRA | 3 |
| 2018 | GOMSF: Graph-Optimization Based Multi-Sensor Fusion for robust UAV Pose estimationabstractAchieving accurate, high-rate pose estimates from proprioceptive and/or exteroceptive measurements is the first step in the development of navigation algorithms for agile mobile robots such as Unmanned Aerial Vehicles (UAVs). In this paper, we propose a decoupled Graph-Optimization based Multi-Sensor Fusion approach (GOMSF) that combines generic 6 Degree-of-Freedom (DoF) visual-inertial odometry poses and 3 DoF globally referenced positions to infer the global 6 DoF pose of the robot in real-time. Our approach casts the fusion as a real-time alignment problem between the local base frame of the visual-inertial odometry and the global base frame. The alignment transformation that relates these coordinate systems is continuously updated by optimizing a sliding window pose graph containing the most recent robot's states. We evaluate the presented pose estimation method on both simulated data and large outdoor experiments using a small UAV that is capable to run our system onboard. Results are compared against different state-of-the-art sensor fusion frameworks, revealing that the proposed approach is substantially more accurate than other decoupled fusion strategies. We also demonstrate comparable results in relation with a finely tuned Extended Kalman Filter that fuses visual, inertial and GPS measurements in a coupled way and show that our approach is generic enough to deal with different input sources in a straightforward manner. Video - https//youtu.be/GIZNSZ2soL8. Ruben Mascaro, Lucas Teixeira, Timo Hinzmann, Roland Siegwart, Margarita Chli |
ICRA | 5 |
| 2017 | Loop-Closure Detection in Urban Scenes for Autonomous Robot NavigationabstractRelocalization is a vital process for autonomous robot navigation, typically running in the background of sequential localization and mapping to detect loops in the robot's trajectory. Such loop-closure detections enable corrections for drift accumulated during the estimation processes and even recovery from complete localization failures. In this work, we present a novel approach loosely integrated with a keyframe-based SLAM system to perform loop-closure detection in urban scenarios for autonomous robot navigation. Generating a mesh of the current robot's surroundings in real-time using monocular and inertial cues, the proposed method estimates the most salient plane in the current view, enabling the creation of the corresponding orthophoto for this plane. Evaluating image similarity on orthophotos forms a much better conditioned problem for relocalization, minimizing effects from viewpoint changes. Employing binary image descriptors and tests on their relative constellation in the image, the proposed approach exhibits robustness also to illumination and situational variations common in real scenes, overall resulting to significant improvement in loop-closure detection performance in urban scenes with respect to the state of the art. Fabiola Maffra, Lucas Teixeira, Zetao Chen, Margarita Chli |
3DV | 4 |
| 2017 | Short-term UAV path-planning with monocular-inertial SLAM in the loopabstractSmall Unmanned Aerial Vehicles (UAVs) are some of the most promising robotic platforms in a variety of applications due to their high mobility. Their restricted computational and payload capabilities, however, translate into significant challenges in automating their navigation. With Simultaneous Localization And Mapping (SLAM) systems recently demonstrated to be employable onboard UAVs, the focus fall on path-planning on the quest of achieving autonomous navigation. With the vast body of path-planning literature often assuming perfect maps or maps known a priori, the biggest challenge lies in dealing with the robustness and accuracy limitations of onboard SLAM in real missions. In this spirit, this paper proposes a path-planning algorithm designed to work in the loop of the SLAM estimation of a monocular-inertial system. This point-to-point planner is demonstrated to navigate in an unknown environment using the incrementally generated SLAM map, while dictating the navigation strategy for preferable acquisition of sensor data for better estimations within SLAM. A thorough evaluation testbed of both simulated and real data is presented, demonstrating the robustness of the proposed pipeline against the state-of-the-art and its dramatically lower computational complexity, revealing its suitability to UAV navigation. Ignacio Alzugaray, Lucas Teixeira, Margarita Chli |
ICRA | 3 |
| 2017 | Autonomous navigation of hexapod robots with vision-based controller adaptationabstractThis work introduces a novel hybrid control architecture for a hexapod platform (Weaver), making it capable of autonomously navigating in uneven terrain. The main contribution stems from the use of vision-based exteroceptive terrain perception to adapt the robot's locomotion parameters. Avoiding computationally expensive path planning for the individual foot tips, the adaptation controller enables the robot to reactively adapt to the surface structure it is moving on. The virtual stiffness, which mainly characterizes the behavior of the legs' impedance controller is adapted according to visually perceived terrain properties. To further improve locomotion, the frequency and height of the robot's stride are similarly adapted. Furthermore, novel methods for terrain characterization and a keyframe based visual-inertial odometry algorithm are combined to generate a spatial map of terrain characteristics. Localization via odometry also allows for autonomous missions on variable terrain by incorporating global navigation and terrain adaptation into one control architecture. Autonomous runs on a testbed with variable terrain types illustrate that adaptive stride and impedance behavior decreases the cost of transport by 30 % compared to a non-adaptive approach and simultaneously increases body stability (up to 88 % on even terrain and by 54 % on uneven terrain). Weaver is able to freely explore outdoor environments as it is completely free of external tethers, as shown in the experiments. Marko Bjelonic, Timon Homberger, Navinda Kottege, Paulo Vinicius Koerich Borges, Margarita Chli, Philipp Beckerle |
ICRA | 5 |
| 2017 | Robust visual-inertial localization with weak GPS priors for repetitive UAV flightsabstractAgile robots, such as small Unmanned Aerial Vehicles (UAVs) can have a great impact on the automation of tasks, such as industrial inspection and maintenance or crop monitoring and fertilization in agriculture. Their deploy-ability, however, relies on the UAV's ability to self-localize with precision and exhibit robustness to common sources of uncertainty in real missions. Here, we propose a new system using the UAV's onboard visual-inertial sensor suite to first build a Reference Map of the UAV's workspace during a piloted reconnaissance flight. In subsequent flights over this area, the proposed framework combines keyframe-based visual-inertial odometry with novel geometric image-based localization, to provide a real-time estimate of the UAV's pose with respect to the Reference Map paving the way towards completely automating repeated navigation in this workspace. The stability of the system is ensured by decoupling the local visual-inertial odometry from the global registration to the Reference Map, while GPS feeds are used as a weak prior for suggesting loop closures. The proposed framework is shown to outperform GPS localization significantly and diminishes drift effects via global image-based alignment for consistently robust performance. Julian Surber, Lucas Teixeira, Margarita Chli |
ICRA | 3 |
| 2017 | Real-time local 3D reconstruction for aerial inspection using superpixel expansionabstractOn the quest of automating the navigation of challenging and promising Robotics platforms such as small Unmanned Aerial Vehicles (UAVs), the community has been increasingly active in developing perception capabilities able to run onboard such platforms in real-time. Despite that vision-based techniques have been at the heart of recent advancements, the realistic employment onboard UAVs is still in its infancy. Inspired by some of the most recent breakthroughs in online dense scene estimation and borrowing fundamental concepts from Computer Vision, in this work we propose a new pipeline for real-time, local scene reconstruction using a single camera for aerial navigation. Aiming for denser scene estimation than traditional feature-based maps with the ability to run onboard a small UAV in real-time, the proposed approach is demonstrated to achieve unprecedented performance producing rich maps of the camera's workspace, timely enough to serve in obstacle avoidance and real-time interaction of a robot with its direct surroundings. Evaluation on benchmarking datasets and on challenging aerial footage captured with a UAV featuring a conventional camera, reveals dramatic speed-ups, as well as denser and more accurate local reconstructions with respect to the state of the art. Lucas Teixeira, Margarita Chli |
ICRA | 2 |
| 2017 | Only look once, mining distinctive landmarks from ConvNet for visual place recognitionabstractRecently, image representations derived from Convolutional Neural Networks (CNNs) have been demonstrated to achieve impressive performance on a wide variety of tasks, including place recognition. In this paper, we take a step deeper into the internal structure of CNNs and propose novel CNN-based image features for place recognition by identifying salient regions and creating their regional representations directly from the convolutional layer activations. A range of experiments is conducted on challenging datasets with varied conditions and viewpoints. These reveal superior precision-recall characteristics and robustness against both viewpoint and appearance variations for the proposed approach over the state of the art. By analyzing the feature encoding process of our approach, we provide insights into what makes an image presentation robust against external variations. Zetao Chen, Fabiola Maffra, Inkyu Sa, Margarita Chli |
IROS | 4 |
| 2016 | Real-time dense surface reconstruction for aerial manipulationabstractWith robotic systems reaching considerable maturity in basic self-localization and environment mapping, new research avenues open up pushing for interaction of a robot with its surroundings for added autonomy. However, the transition from traditionally sparse feature-based maps to dense and accurate scene-estimation imperative for realistic manipulation is not straightforward. Moreover, achieving this level of scene perception in real-time from a computationally constrained and highly shaky and agile platform, such as a small an Unmanned Aerial Vehicle (UAV) is perhaps the most challenging scenario for perception for manipulation. Drawing inspiration from otherwise computationally constraining Computer Vision techniques, we present a system combining visual, inertial and depth information to achieve dense, local scene reconstruction of high precision in real-time. Our evaluation testbed is formed using ground-truth not only in the pose of the sensor-suite, but also the scene reconstruction using a highly accurate laser scanner, offering unprecedented comparisons of scene estimation to ground-truth using real sensor data. Given the lack of any real, ground-truth datasets for environment reconstruction, our V4RL Dense Surface Reconstruction dataset is publicly available. Marco Karrer, Mina Kamel 0001, Roland Siegwart, Margarita Chli |
IROS | 4 |
| 2016 | Real-time mesh-based scene estimation for aerial inspectionabstractWith society and industry pushing for robot-assisted systems to automate cumbersome tasks, such as inspection and maintenance, a vast amount of research effort has been dedicated to relevant technologies. Right at the forefront are small Unmanned Aerial Vehicles (UAVs) equipped with onboard cameras, recently demonstrating that vision-based autonomous flights without reliance on GPS are possible, sparking great interest in a plethora of areas. Current solutions, however, still lack in portability and generality struggling to perform outside the controlled laboratory environment, with onboard robotic perception constituting the biggest impediment. Driven by the need for real-time denser scene estimation, in this work we present a dramatically cheap approach enabling estimation of the immediate surroundings of a UAV using the inertial and visual cues from a single onboard camera. Instead of following the recent trend towards dense scene reconstruction, we trade detail of reconstruction for efficiency of estimation, albeit without compromising accuracy. We also present ETHZ CAB building dataset for aerial inspection. We present results against scene ground truth obtained by a millimetre-precise laser scanner. Lucas Teixeira, Margarita Chli |
IROS | 2 |
| 2015 | Location graphs for visual place recognitionabstractWith the growing demand for deployment of robots in real scenarios, robustness in the perception capabilities for navigation lies at the forefront of research interest, as this forms the backbone of robotic autonomy. Existing place recognition approaches traditionally follow the feature-based bag-of-words paradigm in order to cut down on the richness of information in images. As structural information is typically ignored, such methods suffer from perceptual aliasing and reduced recall, due to the ambiguity of observations. In a bid to boost the robustness of appearance-based place recognition, we consider the world as a continuous constellation of visual words, while keeping track of their covisibility in a graph structure. Locations are queried based on their appearance, and modelled by their corresponding cluster of landmarks from the global covisibility graph, which retains important relational information about landmarks. Complexity is reduced by comparing locations by their graphs of visual words in a simplified manner. Test results show increased recall performance and robustness to noisy observations, compared to state-of-the-art methods. Elena Stumm, Christopher Mei, Simon Lacroix, Margarita Chli |
ICRA | 4 |
| 2014 | People detection and tracking from aerial thermal viewsabstractDetection and tracking of people in visible-light images has been subject to extensive research in the past decades with applications ranging from surveillance to search-and-rescue. Following the growing availability of thermal cameras and the distinctive thermal signature of humans, research effort has been focusing on developing people detection and tracking methodologies applicable to this sensing modality. However, a plethora of challenges arise on the transition from visible-light to thermal images, especially with the recent trend of employing thermal cameras onboard aerial platforms (e.g. in search-and-rescue research) capturing oblique views of the scenery. This paper presents a new, publicly available dataset of annotated thermal image sequences, posing a multitude of challenges for people detection and tracking. Moreover, we propose a new particle filter based framework for tracking people in aerial thermal images. Finally, we evaluate the performance of this pipeline on our dataset, incorporating a selection of relevant, state-of-the-art methods and present a comprehensive discussion of the merits spawning from our study. Jan Portmann, Simon Lynen, Margarita Chli, Roland Siegwart |
ICRA | 3 |
| 2013 | Path planning for motion dependent state estimation on micro aerial vehiclesabstractWith navigation algorithms reaching a certain maturity in the field of mobile robots, the community now focuses on more advanced tasks like path planning towards increased autonomy. While the goal is to efficiently compute a path to a target destination, the uncertainty in the robot's perception cannot be ignored if a realistic path is to be computed. With most state of the art navigation systems providing the uncertainty in motion estimation, here we propose to exploit this information. This leads to a system that can plan safe avoidance of obstacles, and more importantly, it can actively aid navigation by choosing a path that minimizes the uncertainty in the monitored states. Our proposed approach is applicable to systems requiring certain excitations in order to render all their states observable, such as a MAV with visual-inertial based localization. In this work, we propose an approach which takes into account this necessary motion during path planning: by employing Rapidly exploring Random Belief Trees (RRBT), the proposed approach chooses a path to a goal which allows for best estimation of the robot's states, while inherently avoiding motion in unobservable modes. We discuss our findings within the scenario of vision-based aerial navigation as one of the most challenging navigation problem, requiring sufficient excitation to reach full observability. Markus Achtelik, Stephan Weiss 0002, Margarita Chli, Roland Siegwart |
ICRA | 3 |
| 2013 | Inversion based direct position control and trajectory following for micro aerial vehiclesabstractIn this work, we present a powerful, albeit simple position control approach for Micro Aerial Vehicles (MAVs) targeting specifically multicopter systems. Exploiting the differential flatness of four of the six outputs of multicopters, namely position and yaw, we show that the remaining outputs of pitch and roll need not be controlled states, but rather just need to be known. Instead of the common approach of having multiple cascaded control loops (position - velocity - acceleration/attitude - angular rates), the proposed method employs an outer control loop based on dynamic inversion, which directly commands angular rates and thrust. The inner control loop then reduces to a simple proportional controller on the angular rates. As a result, not only does this combination allow for higher bandwidth compared to common control approaches, but also eliminates many mathematical operations (only one trigonometric function is called), speeding up the necessary processing especially on embedded systems. This approach assumes a reliable state estimation framework, which we are able to provide with through previous work. As a result, with this work, we provide the missing elements necessary for a complete approach on autonomous navigation of MAVs. Markus Achtelik, Simon Lynen, Margarita Chli, Roland Siegwart |
IROS | 3 |
| 2013 | A robust and modular multi-sensor fusion approach applied to MAV navigationabstractIt has been long known that fusing information from multiple sensors for robot navigation results in increased robustness and accuracy. However, accurate calibration of the sensor ensemble prior to deployment in the field as well as coping with sensor outages, different measurement rates and delays, render multi-sensor fusion a challenge. As a result, most often, systems do not exploit all the sensor information available in exchange for simplicity. For example, on a mission requiring transition of the robot from indoors to outdoors, it is the norm to ignore the Global Positioning System (GPS) signals which become freely available once outdoors and instead, rely only on sensor feeds (e.g., vision and laser) continuously available throughout the mission. Naturally, this comes at the expense of robustness and accuracy in real deployment. This paper presents a generic framework, dubbed MultiSensor-Fusion Extended Kalman Filter (MSF-EKF), able to process delayed, relative and absolute measurements from a theoretically unlimited number of different sensors and sensor types, while allowing self-calibration of the sensor-suite online. The modularity of MSF-EKF allows seamless handling of additional/lost sensor signals during operation while employing a state buffering scheme augmented with Iterated EKF (IEKF) updates to allow for efficient re-linearization of the prediction to get near optimal linearization points for both absolute and relative state updates. We demonstrate our approach in outdoor navigation experiments using a Micro Aerial Vehicle (MAV) equipped with a GPS receiver as well as visual, inertial, and pressure sensors. Simon Lynen, Markus Achtelik, Stephan Weiss 0002, Margarita Chli, Roland Siegwart |
IROS | 4 |
| 2012 | Versatile distributed pose estimation and sensor self-calibration for an autonomous MAVabstractIn this paper, we present a versatile framework to enable autonomous flights of a Micro Aerial Vehicle (MAV) which has only slow, noisy, delayed and possibly arbitrarily scaled measurements available. Using such measurements directly for position control would be practically impossible as MAVs exhibit great agility in motion. In addition, these measurements often come from a selection of different onboard sensors, hence accurate calibration is crucial to the robustness of the estimation processes. Here, we address these problems using an EKF formulation which fuses these measurements with inertial sensors. We do not only estimate pose and velocity of the MAV, but also estimate sensor biases, scale of the position measurement and self (inter-sensor) calibration in real-time. Furthermore, we show that it is possible to obtain a yaw estimate from position measurements only. We demonstrate that the proposed framework is capable of running entirely onboard a MAV performing state prediction at the rate of 1 kHz. Our results illustrate that this approach is able to handle measurement delays (up to 500ms), noise (std. deviation up to 20 cm) and slow update rates (as low as 1 Hz) while dynamic maneuvers are still possible. We present a detailed quantitative performance evaluation of the real system under the influence of different disturbance parameters and different sensor setups to highlight the versatility of our approach. Stephan Weiss 0002, Markus Achtelik, Margarita Chli, Roland Siegwart |
ICRA | 3 |
| 2012 | Real-time onboard visual-inertial state estimation and self-calibration of MAVs in unknown environmentsabstractThe combination of visual and inertial sensors has proved to be very popular in robot navigation and, in particular, Micro Aerial Vehicle (MAV) navigation due the flexibility in weight, power consumption and low cost it offers. At the same time, coping with the big latency between inertial and visual measurements and processing images in real-time impose great research challenges. Most modern MAV navigation systems avoid to explicitly tackle this by employing a ground station for off-board processing. In this paper, we propose a navigation algorithm for MAVs equipped with a single camera and an Inertial Measurement Unit (IMU) which is able to run onboard and in real-time. The main focus here is on the proposed speed-estimation module which converts the camera into a metric body-speed sensor using IMU data within an EKF framework. We show how this module can be used for full self-calibration of the sensor suite in real-time. The module is then used both during initialization and as a fall-back solution at tracking failures of a keyframe-based VSLAM module. The latter is based on an existing high-performance algorithm, extended such that it achieves scalable 6DoF pose estimation at constant complexity. Fast onboard speed control is ensured by sole reliance on the optical flow of at least two features in two consecutive camera frames and the corresponding IMU readings. Our nonlinear observability analysis and our real experiments demonstrate that this approach can be used to control a MAV in speed, while we also show results of operation at 40Hz on an onboard Atom computer 1.6 GHz. Stephan Weiss 0002, Markus Achtelik, Simon Lynen, Margarita Chli, Roland Siegwart |
ICRA | 4 |
| 2012 | SFly: Swarm of micro flying robotsabstractThe SFly project is an EU-funded project, with the goal to create a swarm of autonomous vision controlled micro aerial vehicles. The mission in mind is that a swarm of MAV's autonomously maps out an unknown environment, computes optimal surveillance positions and places the MAV's there and then locates radio beacons in this environment. The scope of the work includes contributions on multiple different levels ranging from theoretical foundations to hardware design and embedded programming. One of the contributions is the development of a new MAV, a hexacopter, equipped with enough processing power for onboard computer vision. A major contribution is the development of monocular visual SLAM that runs in real-time onboard of the MAV. The visual SLAM results are fused with IMU measurements and are used to stabilize and control the MAV. This enables autonomous flight of the MAV, without the need of a data link to a ground station. Within this scope novel analytical solutions for fusing IMU and vision measurements have been derived. In addition to the realtime local SLAM, an offline dense mapping process has been developed. For this the MAV's are equipped with a payload of a stereo camera system. The dense environment map is used to compute optimal surveillance positions for a swarm of MAV's. For this an optimiziation technique based on cognitive adaptive optimization has been developed. Finally, the MAV's have been equipped with radio transceivers and a method has been developed to locate radio beacons in the observed environment. Markus Achtelik, Michael Achtelik, Yorick Brunet, Margarita Chli, Savvas A. Chatzichristofis, Jean-Dominique Decotignie, Klaus-Michael Doth, Friedrich Fraundorfer, Laurent Kneip, Daniel Gurdan, Lionel Heng, Elias B. Kosmatopoulos, Lefteris Doitsidis, Gim Hee Lee, Simon Lynen, Agostino Martinelli, Lorenz Meier, Marc Pollefeys, Damien Piguet, Alessandro Renzaglia, Davide Scaramuzza 0001, Roland Siegwart, Jan Stumpf, Petri Tanskanen, Chiara Troiani, Stephan Weiss 0002 |
IROS | 4 |
| 2012 | Visual-inertial SLAM for a small helicopter in large outdoor environmentsabstractIn this video, we present our latest results towards fully autonomous flights with a small helicopter. Using a monocular camera as the only exteroceptive sensor, we fuse inertial measurements to achieve a self-calibrating power-on-and-go system, able to perform autonomous flights in previously unknown, large, outdoor spaces. Our framework achieves Simultaneous Localization And Mapping (SLAM) with previously unseen robustness in onboard aerial navigation for small platforms with natural restrictions on weight and computational power. We demonstrate successful operation in flights with altitude between 0.2-70 m, trajectories with 350 m length, as well as dynamic maneuvers with track speed of 2 m/s. All flights shown are performed autonomously using vision in the loop, with only high-level waypoints given as directions. Markus Achtelik, Simon Lynen, Stephan Weiss 0002, Laurent Kneip, Margarita Chli, Roland Siegwart |
IROS | 5 |
| 2011 | Robust Real-Time Visual Odometry with a Single Camera and an IMUabstractThe increasing demand for real-time high-precision Visual Odometry systems as part of navigation and localization tasks has recently been driving research towards more versatile and scalable solutions. In this paper, we present a novel framework for combining the merits of inertial and visual data from a monocular camera to accumulate estimates of local motion incrementally and reliably reconstruct the trajectory traversed. We demonstrate the robustness and efficiency of our methodology in a scenario with challenging camera dynamics, and present a comprehensive evaluation against ground-truth data. 1 Laurent Kneip, Margarita Chli, Roland Siegwart |
BMVC | 2 |
| 2011 | BRISK: Binary Robust invariant scalable keypointsabstractEffective and efficient generation of keypoints from an image is a well-studied problem in the literature and forms the basis of numerous Computer Vision applications. Established leaders in the field are the SIFT and SURF algorithms which exhibit great performance under a variety of image transformations, with SURF in particular considered as the most computationally efficient amongst the high-performance methods to date. In this paper we propose BRISK1, a novel method for keypoint detection, description and matching. A comprehensive evaluation on benchmark datasets reveals BRISK's adaptive, high quality performance as in state-of-the-art algorithms, albeit at a dramatically lower computational cost (an order of magnitude faster than SURF in cases). The key to speed lies in the application of a novel scale-space FAST-based detector in combination with the assembly of a bit-string descriptor from intensity comparisons retrieved by dedicated sampling of each keypoint neighborhood. Stefan Leutenegger, Margarita Chli, Roland Siegwart |
ICCV | 2 |
| 2011 | Collaborative stereoabstractIn this paper, we propose a method to recover the relative pose of two robots in absolute scale and in real-time using one monocular camera on each robot. We achieve this by fusing measurements from the onboard inertial sensors on each platform with information obtained from feature correspondences between the two cameras using an Extended Kalman Filter (EKF). This forms a flexible stereo rig, providing the ability to treat the two robots as one single dynamic sensor, which can adapt to the environment and thus improve environmental mapping, obstacle avoidance and navigation. We demonstrate the power of this approach on both simulation and real datasets, employing two micro aerial vehicles (MAVs) to illustrate successful operation over general 3D motion. Markus Achtelik, Stephan Weiss 0002, Margarita Chli, Frank Dellaert, Roland Siegwart |
IROS | 3 |
| 2010 | Scalable active matchingabstractIn matching tasks in computer vision, and particularly in real-time tracking from video, there are generally strong priors available on absolute and relative correspondence locations thanks to motion and scene models. While these priors are often partially used post-hoc to resolve matching consensus in algorithms like RANSAC, it was recently shown that fully integrating them in an `Active Matching' (AM) approach permits efficient guided image processing with rigorous decisions guided by Information Theory. AM's weakness was that the overhead induced by intermediate Bayesian updates required meant poor scaling to cases where many correspondences were sought. In this paper we show that relaxation of the rigid probabilistic model of AM, where every feature measurement directly affects the prediction of every other, permits dramatically more scalable operation without affecting accuracy. We take a general graph-theoretic view of the structure of prior information in matching to sparsify and approximate the interconnections. We demonstrate the performance of two variations, CLAM and SubAM, in the context of sequential camera tracking. These algorithms are highly competitive with other techniques at matching hundreds of features per frame while retaining great intuitive appeal and the full probabilistic capability to digest prior information. Ankur Handa, Margarita Chli, Hauke Strasdat, Andrew J. Davison |
CVPR | 2 |
| 2009 | Automatically and efficiently inferring the hierarchical structure of visual mapsabstractIn Simultaneous Localisation and Mapping (SLAM), it is well known that probabilistic filtering approaches which aim to estimate the robot and map state sequentially suffer from poor computational scaling to large map sizes. Various authors have demonstrated that this problem can be mitigated by approximations which treat estimates of features in different parts of a map as conditionally independent, allowing them to be processed separately. When it comes to the choice of how to divide a large map into such ‘submaps’, straightforward heuristics may be sufficient in maps built using sensors such as laser range-finders with limited range, where a regular grid of submap boundaries performs well. With visual sensing, however, the ideal division of submaps is less clear, since a camera has potentially unlimited range and will often observe spatially distant parts of a scene simultaneously. In this paper we present an efficient and generic method for automatically determining a suitable submap division for SLAM maps, and apply this to visual maps built with a single agile camera. We use the mutual information between predicted measurements of features as an absolute measure of correlation, and cluster highly correlated features into groups. Via tree factorisation, we are able to determine not just a single level submap division but a powerful fully hierarchical correlation and clustering structure. Our analysis and experiments reveal particularly interesting structure in visual maps and give pointers to more efficient approximate visual SLAM algorithms. Margarita Chli, Andrew J. Davison |
ICRA | 1 |
| 2008 | Active Matching
Margarita Chli, Andrew J. Davison |
ECCV (1) | 1 |