Viorela Ila

dblp:09/789 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
5since 2021 · last 2026
0000-0002-8137-0833ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 4 first-author · 4 since 2021Systems, architecture and hardware · 14 · 4 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 DynoSAM: Open-Source Smoothing and Mapping Framework for Dynamic SLAM
abstract
Traditional Visual Simultaneous Localization and Mapping systems focus solely on static scene structures, overlooking dynamic elements in the environment. Although effective for accurate visual odometry in complex scenarios, these methods discard crucial information about moving objects. By incorporating this information into a Dynamic SLAM framework, the motion of dynamic entities can be estimated, enhancing navigation whilst ensuring accurate localization. However, the fundamental formulation of Dynamic SLAM remains an open challenge, with no consensus on the optimal approach for accurate motion estimation within a SLAM pipeline. Therefore, we developedDynoSAM, an open-source framework for Dynamic Objects SLAM that enables the efficient implementation, testing, and comparison of various Dynamic SLAM optimization formulations. We further propose a novel formulation that encodes rigid-body motion model in object pose estimation as well as an error metric agnostic to object frame definition.DynoSAMintegrates static and dynamic measurements into a unified optimization problem solved using factor graphs, simultaneously estimating camera poses, static scene, object motion or poses, and object structures. We evaluateDynoSAMacross diverse simulated and real-world datasets, achieving state-of-the-art motion estimation in indoor and outdoor environments, with substantial improvements over existing systems. Additionally, we demonstrateDynoSAM's contributions to downstream applications, including 3D reconstruction of dynamic scenes and trajectory prediction, thereby showcasing potential for advancing dynamic object-aware SLAM systems. Code is open-sourced athttps://github.com/ACFR-RPG/DynOSAM
Jesse Morris, Yiduo Wang 0001, Mikolaj Kliniewski, Viorela Ila
IEEE Trans. Robotics4
2025 TB-HSU: Hierarchical 3D Scene Understanding with Contextual Affordances
abstract
The concept of function and affordance is a critical aspect of 3D scene understanding and supports task-oriented objectives. In this work, we develop a model that learns to structure and vary functional affordance across a 3D hierarchical scene graph representing the spatial organization of a scene. The varying functional affordance is designed to integrate with the varying spatial context of the graph. More specifically, we develop an algorithm that learns to construct a 3D hierarchical scene graph (3DHSG) that captures the spatial organization of the scene. Starting from segmented object point clouds and object semantic labels, we develop a 3DHSG with a top node that identifies the room label, child nodes that define local spatial regions inside the room with region-specific affordances, and grand-child nodes indicating object locations and object-specific affordances. To support this work, we create a custom 3DHSG dataset that provides ground truth data for local spatial regions with region-specific affordances and also object-specific affordances for each object. We employ a Transformer Based Hierarchical Scene Understanding (TB-HSU) model to learn the 3DHSG. We use a multi-task learning framework that learns both room classification and learns to define spatial regions within the room with region-specific affordances. Our work improves on the performance of state-of-the-art baseline models and shows one approach for applying transformer models to 3D scene understanding and the generation of 3DHSGs that capture the spatial organization of a room. The code and dataset are publicly available.
Wenting Xu, Viorela Ila, Luping Zhou, Craig T. Jin
AAAI2
2025 DynORecon: Dynamic Object Reconstruction for Navigation
abstract
This paper presents DynORecon, a Dynamic Object Reconstruction system that leverages the information provided by Dynamic SLAM to simultaneously generate a volumetric map of observed moving entities while estimating free space to support navigation. By capitalising on the motion estimations provided by Dynamic SLAM, DynORecon continuously refines the representation of dynamic objects to eliminate residual artefacts from past observations and incrementally reconstructs each object, seamlessly integrating new observations to capture previously unseen structures. Our system is highly efficient (~20 FPS) and produces accurate (~10 cm) object reconstructions using simulated and real-world outdoor datasets.
Yiduo Wang 0001, Jesse Morris, Teresa Vidal-Calleja, Viorela Ila
ICRA5
2024 The Importance of Coordinate Frames in Dynamic SLAM
abstract
Most Simultaneous localisation and mapping (SLAM) systems have traditionally assumed a static world, which does not align with real-world scenarios. To enable robots to safely navigate and plan in dynamic environments, it is essential to employ representations capable of handling moving objects. Dynamic SLAM is an emerging field in SLAM research as it improves the overall system accuracy while providing additional estimation of object motions. State-of-the-art literature informs two main formulations for Dynamic SLAM, representing dynamic object points in either the world or object coordinate frame. While expressing object points in their local reference frame may seem intuitive, it does not necessarily lead to the most accurate and robust solutions. This paper conducts and presents a thorough analysis of various Dynamic SLAM formulations, identifying the best approach to address the problem. To this end, we introduce a front-end agnostic framework using GTSAM [1] that can be used to evaluate various Dynamic SLAM formulations.1
Jesse Morris, Yiduo Wang 0001, Viorela Ila
ICRA3
2023 Principled ICP Covariance Modelling in Perceptually Degraded Environments for the EELS Mission Concept
abstract
The Exobiology Extant Life Surveyor (EELS) is a snake-like mobile instruments platform under development at Jet Propulsion Laboratory (JPL) for a mission concept to find evidence of life on Saturn's sixth largest moon, Enceladus. To conduct a life surveying mission there, the EELS platform must first traverse an unknown icy surface terrain before undertaking a controlled descent into a cryovolcanic vent. The remoteness of Enceladus and the icy nature of its terrain demands a level of autonomy in navigation significantly higher than previous rover missions. The perception system onboard EELS must be highly resilient to perceptually-degraded environments such as flat, open ice fields, icy plumes, and repeating geometries in vents. EELS' perception system is implemented as a multi-sensor Simultaneous Localisation And Mapping (SLAM) solution called SERPENT. State Estimation through Robust Perception in Extreme and Novel Terrains (SERPENT) estimates the robot trajectory and maintains a map database, from which dense global or local maps can be obtained on demand for downstream planning algorithms. This system opts to incorporate measurements from many sensor modalities (laser scans, images, IMU, altimeter, etc.), solving the SLAM problem through joint optimisation, and thus requires that the contribution of each sensor be balanced through careful modelling of their uncertainties. With a specific focus on Light Detection And Ranging (LiDAR) in this context, this paper proposes a principled approach to model the covariances of point-to-plane Iterative Closest Point (ICP). It performs a rigorous comparative analysis of new and existing covariance models, and is the first time some of these have been tested within a complete SLAM pipeline. These models are evaluated on perceptually challenging datasets collected in glacial environments by the EELS sensor suite (see Figures 1, 2). SERPENT is open-sourced at https://github.com/jpl-eels/serpent.
William Talbot, Jeremy Nash, Michael Paton, Eric Ambrose, Brandon Metz, Rohan Thakker, Rachel Etheredge, Masahiro Ono, Viorela Ila
IROS9
2020 Dynamic SLAM: The Need For Speed
abstract
The static world assumption is standard in most simultaneous localisation and mapping (SLAM) algorithms. Increased deployment of autonomous systems to unstructured dynamic environments is driving a need to identify moving objects and estimate their velocity in real-time. Most existing SLAM based approaches rely on a database of 3D models of objects or impose significant motion constraints. In this paper, we propose a new feature-based, model-free, object-aware dynamic SLAM algorithm that exploits semantic segmentation to allow estimation of motion of rigid objects in a scene without the need to estimate the object poses or have any prior knowledge of their 3D models. The algorithm generates a map of dynamic and static structure and has the ability to extract velocities of rigid moving objects in the scene. Its performance is demonstrated on simulated, synthetic and real-world datasets.
Mina Henein, Robert E. Mahony, Viorela Ila
ICRA4
2020 Robust Ego and Object 6-DoF Motion Estimation and Tracking
abstract
The problem of tracking self-motion as well as motion of objects in the scene using information from a camera is known as multi-body visual odometry and is a challenging task. This paper proposes a robust solution to achieve accurate estimation and consistent track-ability for dynamic multi-body visual odometry. A compact and effective framework is proposed leveraging recent advances in semantic instance-level segmentation and accurate optical flow estimation. A novel formulation, jointly optimizing SE(3) motion and optical flow is introduced that improves the quality of the tracked points and the motion estimation accuracy. The proposed approach is evaluated on the virtual KITTI Dataset and tested on the real KITTI Dataset, demonstrating its applicability to autonomous driving applications. For the benefit of the community, we make the source code public†.
Mina Henein, Robert E. Mahony, Viorela Ila
IROS4
2018 Calibrating Light-Field Cameras Using Plenoptic Disc Features
abstract
This paper proposes a new method for estimating calibration parameters of plenoptic cameras by minimizing the nonlinear plenoptic reprojection error. Novel plenoptic feature types are proposed as data for the calibration method. These plenoptic disc features are in a natural one-to-one correspondence with physical points in front of the camera. We exploit the intrinsic geometry of plenoptic cameras in a novel projection model that relates the plenoptic disc features to physical points. The resulting calibration quality, as quantified by mean reprojection error and 3D reconstruction error, outperforms recently published results.
Sean G. P. O'Brien, Jochen Trumpf, Viorela Ila, Robert E. Mahony
3DV3
2017 Fast Incremental Bundle Adjustment with Covariance Recovery
abstract
Efficient algorithms exist to obtain a sparse 3D representation of the environment. Bundle adjustment (BA) and structure from motion (SFM) are techniques used to estimate both the camera poses and the set of sparse points in the environment. Many applications require such reconstruction to be performed online, while acquiring the data, and produce an updated result every step. Furthermore, using active feedback about the quality of the reconstruction can help selecting the best views to increase the accuracy as well as to maintain a reasonable size of the collected data. This paper provides novel and efficient solutions to solving the associated NLS incrementally, and to compute not only the optimal solution, but also the associated uncertainty. The proposed technique highly increases the efficiency of the incremental BA solver for long camera trajectory applications, and provides extremely fast covariance recovery.
Viorela Ila, Lukás Polok, Marek Solony, Klemen Istenic
3DV1
2017 Exploring the effect of meta-structural information on the global consistency of SLAM
abstract
Accurate online estimation of the environment structure simultaneously with the robot pose is a key capability for autonomous robotic vehicles. Classical simultaneous localization and mapping (SLAM) algorithms make no assumptions about the configuration of the points in the environment, however, real world scenes have significant structure (ground planes, buildings, walls, ceilings, etc.) that can be exploited. In this paper, we introduce meta-structural information associated with geometric primitives into the estimation problem and analyze their effect on the global structural consistency of the resulting map. Although we only consider the effect of adding planar and orthogonality information for the estimation of 3D points in a Manhattan-like world, this framework can be extended to any type of geometric, kinematic, dynamic or even semantic information. We evaluate our approach on a city-like simulated environment. We highlight the advantages of the proposed solution over SLAM formulation considering no prior knowledge about the configuration of 3D points in the environment.
Mina Henein, Montiel Abello, Viorela Ila, Robert E. Mahony
IROS3
2016 Big Data Analysis for Media Production
abstract
A typical high-end film production generates several terabytes of data per day, either as footage from multiple cameras or as background information regarding the set (laser scans, spherical captures, etc). This paper presents solutions to improve the integration of the multiple data sources, and understand their quality and content, which are useful both to support creative decisions on-set (or near it) and enhance the postproduction process. The main cinema specific contributions, tested on a multisource production dataset made publicly available for research purposes, are the monitoring and quality assurance of multicamera set-ups, multisource registration and acceleration of 3-D reconstruction, anthropocentric visual analysis techniques for semantic content annotation, and integrated 2-D–3-D web visualization tools. We discuss as well improvements carried out in basic techniques for acceleration, clustering and visualization, which were necessary to deal with the very large multisource data, and can be applied to other big data problems in diverse application fields.
Josep Blat, Alun Evans, Hansung Kim 0001, Evren Imre, Lukás Polok, Viorela Ila, Nikos Nikolaidis 0001, Pavel Zemcík, Anastasios Tefas, Pavel Smrz, Adrian Hilton 0001, Ioannis Pitas
Proc. IEEE6
2015 Fast covariance recovery in incremental nonlinear least square solvers
abstract
Many estimation problems in robotics rely on efficiently solving nonlinear least squares (NLS). For example, it is well known that the simultaneous localisation and mapping (SLAM) problem can be formulated as a maximum likelihood estimation (MLE) and solved using NLS, yielding a mean state vector. However, for many applications recovering only the mean vector is not enough. Data association, active decisions, next best view, are only few of the applications that require fast state covariance recovery. The problem is not simple since, in general, the covariance is obtained by inverting the system matrix and the result is dense. The main contribution of this paper is a novel algorithm for fast incremental covariance update, complemented by a highly efficient implementation of the covariance recovery. This combination yields to two orders of magnitude reduction in computation time, compared to the other state of the art solutions. The proposed algorithm is applicable to any NLS solver implementation, and does not depend on incremental strategies described in our previous papers, which are not a subject of this paper.
Viorela Ila, Lukás Polok, Marek Solony, Pavel Smrz, Pavel Zemcík
ICRA1
2013 Efficient implementation for block matrix operations for nonlinear least squares problems in robotic applications
abstract
A large number of robotic, computer vision and computer graphics applications rely on efficiently solving the associated sparse linear systems. Simultaneous localization and mapping (SLAM), structure from motion (SfM), non-rigid shape recovery, and elastodynamic simulations are only few examples in this direction. In general, these problems are nonlinear and the solution can be approximated by incrementally solving a series of linearized problems. In some applications, the size of the system considerably affects the performance, especially when the sparsity is low. This paper exploits the block structure of such problems and offers very efficient solutions to manipulate block matrices within iterative nonlinear solvers. The resulting method considerably speeds-up the execution of the implementation of the nonlinear optimization problem. In this work, in particular, we focus our effort on testing the method on SLAM applications, but the applicability of the technique remains general. Our implementation outperforms the state of the art SLAM implementations on all tested datasets. In incremental mode, where a larger portion of time is spent in updating the system, our implementation is on average two times faster than the others.
Lukás Polok, Marek Solony, Viorela Ila, Pavel Smrz, Pavel Zemcík
ICRA3
2011 iSAM2: Incremental smoothing and mapping with fluid relinearization and incremental variable reordering
abstract
We present iSAM2, a fully incremental, graph-based version of incremental smoothing and mapping (iSAM). iSAM2 is based on a novel graphical model-based interpretation of incremental sparse matrix factorization methods, afforded by the recently introduced Bayes tree data structure. The original iSAM algorithm incrementally maintains the square root information matrix by applying matrix factorization updates. We analyze the matrix updates as simple editing operations on the Bayes tree and the conditional densities represented by its cliques. Based on that insight, we present a new method to incrementally change the variable ordering which has a large effect on efficiency. The efficiency and accuracy of the new method is based on fluid relinearization, the concept of selectively relinearizing variables as needed. This allows us to obtain a fully incremental algorithm without any need for periodic batch steps. We analyze the properties of the resulting algorithm in detail, and show on various real and simulated datasets that the iSAM2 algorithm compares favorably with other recent mapping algorithms in both quality and efficiency.
Michael Kaess, Hordur Johannsson, Richard Roberts 0001, Viorela Ila, John J. Leonard, Frank Dellaert
ICRA4
2010 3D reconstruction of underwater structures
abstract
Environmental change is a growing international concern, calling for the regular monitoring, studying and preserving of detailed information about the evolution of underwater ecosystems. For example, fragile coral reefs are exposed to various sources of hazards and potential destruction, and need close observation. Computer vision offers promising technologies to build 3D models of an environment from two-dimensional images. The state of the art techniques have enabled high-quality digital reconstruction of large-scale structures, e.g., buildings and urban environments, but only sparse representations or dense reconstruction of small objects have been obtained from underwater video and still imagery. The application of standard 3D reconstruction methods to challenging underwater environments typically produces unsatisfactory results. Accurate, full camera trajectories are needed to serve as the basis for dense 3D reconstruction. A highly accurate sparse 3D reconstruction is the ideal foundation on which to base subsequent dense reconstruction algorithms. In our application the models are constructed from synchronized high definition videos collected using a wide baseline stereo rig. The rig can be hand-held, attached to a boat, or even to an autonomous underwater vehicle. We solve this problem by employing a smoothing and mapping toolkit developed in our lab specifically for this type of application. The result of our technique is a highly accurate sparse 3D reconstruction of underwater structures such as corals.
Chris Beall, Brian Lawrence, Viorela Ila, Frank Dellaert
IROS3
2010 Subgraph-preconditioned conjugate gradients for large scale SLAM
abstract
In this paper we propose an efficient preconditioned conjugate gradients (PCG) approach to solving large-scale SLAM problems. While direct methods, popular in the literature, exhibit quadratic convergence and can be quite efficient for sparse problems, they typically require a lot of storage and efficient elimination orderings to be found. In contrast, iterative optimization methods only require access to the gradient and have a small memory footprint, but can suffer from poor convergence. Our new method, subgraph preconditioning, is obtained by re-interpreting the method of conjugate gradients in terms of the graphical model representation of the SLAM problem. The main idea is to combine the advantages of direct and iterative methods, by identifying a sub-problem that can be easily solved using direct methods, and solving for the remaining part using PCG. The easy sub-problems correspond to a spanning tree, a planar subgraph, or any other substructure that can be efficiently solved. As such, our approach provides new insights into the performance of state of the art iterative SLAM methods based on re-parameterized stochastic gradient descent. The efficiency of our new algorithm is illustrated on large datasets, both simulated and real.
Frank Dellaert, Justin Carlson, Viorela Ila, Kai Ni 0001, Charles E. Thorpe
IROS3
2010 The Bayes Tree: An Algorithmic Foundation for Probabilistic Robot Mapping
Michael Kaess, Viorela Ila, Richard Roberts 0001, Frank Dellaert
WAFR2
2010 Information-Based Compact Pose SLAM
abstract
Pose SLAM is the variant of simultaneous localization and map building (SLAM) is the variant of SLAM, in which only the robot trajectory is estimated and where landmarks are only used to produce relative constraints between robot poses. To reduce the computational cost of the information filter form of Pose SLAM and, at the same time, to delay inconsistency as much as possible, we introduce an approach that takes into account only highly informative loop-closure links and nonredundant poses. This approach includes constant time procedures to compute the distance between poses, the expected information gain for each potential link, and the exact marginal covariances while moving in open loop, as well as a procedure to recover the state after a loop closure that, in practical situations, scales linearly in terms of both time and memory. Using these procedures, the robot operates most of the time in open loop, and the cost of the loop closure is amortized over long trajectories. This way, the computational bottleneck shifts to data association, which is the search over the set of previously visited poses to determine good candidates for sensor registration. To speed up data association, we introduce a method to search for neighboring poses whose complexity ranges from logarithmic in the usual case to linear in degenerate situations. The method is based on organizing the pose information in a balanced tree whose internal levels are defined using interval arithmetic. The proposed Pose-SLAM approach is validated through simulations, real mapping sessions, and experiments using standard SLAM data sets.
Viorela Ila, Josep M. Porta, Juan Andrade-Cetto
IEEE Trans. Robotics1
2009 Reduced state representation in delayed-state SLAM
abstract
This paper introduces an approach that reduces the size of the state and maximizes the sparsity of the information matrix in exactly sparse delayed-state SLAM. We propose constant time procedures to measure the distance between a given pair of poses, the mutual information gain for a given candidate link, and the joint marginals required for both measures. Using these measures, we can readily identify non redundant poses and highly informative links and use only those to augment and to update the state, respectively. The result is a delayed-state SLAM system that reduces both the use of memory and the execution time and that delays filter inconsistency by reducing the number of linearization introduced when adding new loop closure links. We evaluate the advantage of the proposed approach using simulations and data sets collected with real robots.
Viorela Ila, Josep M. Porta, Juan Andrade-Cetto
IROS1
2007 Vision-based loop closing for delayed state robot mapping
abstract
This paper shows results on outdoor vision-based loop closing for simultaneous localization and mapping. Our experiments show that for loops of over 50 m, the pose estimates maintained with a delayed-state extended information filter are consistent enough to guarantee assertion of vision- based pose constraints for loop closure, provided no necessary information links are added to the estimator. The technique computes relative pose constraints via a robust least squares minimization of 3D point correspondences, which are in turn obtained from the matching of SIFT features over candidate image pairs. We propose a loop closure test that checks both for closeness of means and for highly informative updates at the same time.
Viorela Ila, Juan Andrade-Cetto, Rafael Valencia, Alberto Sanfeliu
IROS1
2005 Interest point characterisation through textural analysis for rejection of bad correspondences
Viorela Ila, Rafael García, Xavier Cufí, Joan Batlle
Pattern Recognit. Lett.1
2004 FPGA Implementation of a Vision-Based Motion Estimation Algorithm for an Underwater Robot
Viorela Ila, Rafael García, François Charot, Joan Batlle
FPL1