VLDB 2026 Research / reviewers in the wild / expert
Matías Mattamala
dblp:175/0277
· DBLP profile ↗
15ranked-venue papers
3as first author
11since 2021 · last 2025
0000-0001-6128-7808ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 15 · 3 first-author · 11 since 2021Systems, architecture and hardware · 10 · 1 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Finite-State Controller Based Offline Solver for Deterministic POMDPsabstractDeterministic partially observable Markov decision processes (DetPOMDPs) often arise in planning problems where the agent is uncertain about its environmental state but can act and observe deterministically. In this paper, we propose DetMCVI, an adaptation of the Monte Carlo Value Iteration (MCVI) algorithm for DetPOMDPs, which builds policies in the form of finite-state controllers (FSCs). DetMCVI solves large problems with a high success rate, outperforming existing baselines for DetPOMDPs. We also verify the performance of the algorithm in a real-world mobile robot forest mapping scenario. Alex Schutz, Yang You 0003, Matías Mattamala, Ipek Caliskanelli, Bruno Lacerda, Nick Hawes |
IJCAI | 3 |
| 2025 | OpenLex3D: A Tiered Benchmark for Open-Vocabulary 3D Scene Representationsabstract3D scene understanding has been transformed by open-vocabulary language models that enable interaction via natural language. However, at present the evaluation of these representations is limited to datasets with closed-set semantics that do not capture the richness of language. This work presents OpenLex3D, a dedicated benchmark for evaluating 3D open-vocabulary scene representations. OpenLex3D provides entirely new label annotations for scenes from Replica, ScanNet++, and HM3D, which capture real-world linguistic variability by introducing synonymical object categories and additional nuanced descriptions. Our label sets provide 13 times more labels per scene than the original datasets. By introducing an open-set 3D semantic segmentation task and an object retrieval task, we evaluate various existing 3D open-vocabulary methods on OpenLex3D, showcasing failure cases, and avenues for improvement. Our experiments provide insights on feature precision, segmentation, and downstream capabilities. The benchmark is publicly available at: https://openlex3d.github.io/. Christina Kassab, Sacha Morin, Martin Büchner, Matías Mattamala, Kumaraditya Gupta, Abhinav Valada, Liam Paull, Maurice Fallon |
NeurIPS | 4 |
| 2024 | Language-EXtended Indoor SLAM (LEXIS): A Versatile System for Real-time Visual Scene UnderstandingabstractVersatile and adaptive semantic understanding would enable autonomous systems to comprehend and interact with their surroundings. Existing fixed-class models limit the adaptability of indoor mobile and assistive autonomous systems. In this work, we introduce LEXIS, a real-time indoor Simultaneous Localization and Mapping (SLAM) system that harnesses the open-vocabulary nature of Large Language Models (LLMs) to create a unified approach to scene understanding and place recognition. The approach first builds a topological SLAM graph of the environment (using visual-inertial odometry) and embeds Contrastive Language-Image Pretraining (CLIP) features in the graph nodes. We use this representation for flexible room classification and segmentation, serving as a basis for room-centric place recognition. This allows loop closure searches to be directed towards semantically relevant places. Our proposed system is evaluated using both public, simulated data and real-world data, covering office and home environments. It successfully categorizes rooms with varying layouts and dimensions and outperforms the state-of-the-art (SOTA). For place recognition and trajectory estimation tasks we achieve equivalent performance to the SOTA, all also utilizing the same pre-trained model. Lastly, we demonstrate the system’s potential for planning. Video at: https://youtu.be/gRqF3euDfX8 Christina Kassab, Matías Mattamala, Lintong Zhang, Maurice Fallon |
ICRA | 2 |
| 2024 | Tree Instance Segmentation and Traits Estimation for Forestry Environments Exploiting LiDAR Data Collected by Mobile RobotsabstractForests play a crucial role in our ecosystems, functioning as carbon sinks, climate stabilizers, biodiversity hubs, and sources of wood. By the very nature of their scale, monitoring and maintaining forests is a challenging task. Robotics in forestry can have the potential for substantial automation toward efficient and sustainable foresting practices. In this paper, we address the problem of automatically producing a forest inventory by exploiting LiDAR data collected by a mobile platform. To construct an inventory, we first extract tree instances from point clouds. Then, we process each instance to extract forestry inventory information. Our approach provides the per-tree geometric trait of "diameter at breast height" together with the individual tree locations in a plot. We validate our results against manual measurements collected by foresters during field trials. Our experiments show strong segmentation and tree trait estimation performance, underlining the potential for automating forestry services. Results furthermore show a superior performance compared to the popular baseline methods used in this domain. Meher V. R. Malladi, Tiziano Guadagnino, Luca Lobefaro, Matías Mattamala, Holger Griess, Janine Schweier, Nived Chebrolu, Maurice Fallon, Jens Behley, Cyrill Stachniss |
ICRA | 4 |
| 2024 | SiLVR: Scalable Lidar-Visual Reconstruction with Neural Radiance Fields for Robotic InspectionabstractWe present a neural-field-based large-scale reconstruction system that fuses lidar and vision data to generate high-quality reconstructions that are geometrically accurate and capture photo-realistic textures. This system adapts the state-of-the-art neural radiance field (NeRF) representation to also incorporate lidar data which adds strong geometric constraints on the depth and surface normals. We exploit the trajectory from a real-time lidar SLAM system to bootstrap a Structure-from-Motion (SfM) procedure to both significantly reduce the computation time and to provide metric scale which is crucial for lidar depth loss. We use submapping to scale the system to large-scale environments captured over long trajectories. We demonstrate the reconstruction system with data from a multi-camera, lidar sensor suite onboard a legged robot, hand-held while scanning building scenes for 600 metres, and onboard an aerial robot surveying a multi-storey mock disaster site-building. Website: https://ori-drs.github.io/projects/silvr/ Yifu Tao, Yash Bhalgat, Lanke Frank Tarimo Fu, Matías Mattamala, Nived Chebrolu, Maurice Fallon |
ICRA | 4 |
| 2024 | Resilient Legged Local Navigation: Learning to Traverse with Compromised Perception End-to-EndabstractAutonomous robots must navigate reliably in unknown environments even under compromised exteroceptive perception, or perception failures. Such failures often occur when harsh environments lead to degraded sensing, or when the perception algorithm misinterprets the scene due to limited generalization. In this paper, we model perception failures as invisible obstacles and pits, and train a reinforcement learning (RL) based local navigation policy to guide our legged robot. Unlike previous works relying on heuristics and anomaly detection to update navigational information, we train our navigation policy to reconstruct the environment information in the latent space from corrupted perception and react to perception failures end-to-end. To this end, we incorporate both proprioception and exteroception into our policy inputs, thereby enabling the policy to sense collisions on different body parts and pits, prompting corresponding reactions. We validate our approach in simulation and on the real quadruped robot ANYmal running in real-time (<10ms CPU inference). In a quantitative comparison with existing heuristic-based locally reactive planners, our policy increases the success rate over 30% when facing perception failures. Project Page: https://bit.ly/45NBTuh. Jonas Frey, Nikita Rudin, Matías Mattamala, Cesar Dario Cadena Lerma, Marco Hutter 0001 |
ICRA | 5 |
| 2024 | Markerless Aerial-Terrestrial Co-Registration of Forest Point Clouds using a Deformable Pose GraphabstractFor biodiversity and forestry applications, end-users desire maps of forests that are fully detailed—from the forest floor to the canopy. Terrestrial laser scanning and aerial laser scanning are accurate and increasingly mature methods for scanning the forest. However, individually they are not able to estimate attributes such as tree height, trunk diameter and canopy density due to the inherent differences in their field-of-view and mapping processes. In this work, we present a pipeline that can automatically generate a single joint terrestrial and aerial forest reconstruction. The novelty of the approach is a marker-free registration pipeline, which estimates a set of relative transformation constraints between the aerial cloud and terrestrial sub-clouds without requiring any co-registration reflective markers to be physically placed in the scene. Our method then uses these constraints in a pose graph formulation, which enables us to finely align the respective clouds while respecting spatial constraints introduced by the terrestrial SLAM scanning process. We demonstrate that our approach can produce a fine-grained and complete reconstruction of large-scale natural environments, enabling multi-platform data capture for forestry applications without requiring external infrastructure. Benoît Casseau, Nived Chebrolu, Matías Mattamala, Leonard Freißmuth, Maurice Fallon |
IROS | 3 |
| 2024 | Online Tree Reconstruction and Forest Inventory on a Mobile Robotic SystemabstractTerrestrial laser scanning (TLS) is the standard technique used to create accurate point clouds for digital forest inventories. However, the measurement process is demanding, requiring up to two days per hectare for data collection, significant data storage, as well as resource-heavy post-processing of 3D data. In this work, we present a real-time mapping and analysis system that enables online generation of forest inventories using mobile laser scanners that can be mounted e.g. on mobile robots. Given incrementally created and locally accurate submaps—data payloads—our approach extracts tree candidates using a custom, Voronoi-inspired clustering algorithm. Tree candidates are reconstructed using an algorithm based on the Hough transform, which enables robust modeling of the tree stem. Further, we explicitly incorporate the incremental nature of the data collection by consistently updating the database using a pose graph LiDAR SLAM system. This enables us to refine our estimates of the tree traits if an area is revisited later during a mission. We demonstrate competitive accuracy to TLS or manual measurements using laser scanners that we mounted on backpacks or mobile robots operating in conifer, broad-leaf and mixed forests. Our results achieve RMSE of 1.93 cm, a bias of 0.65 cm and a standard deviation of 1.81 cm (averaged across these sequences)—with no post-processing required after the mission is complete. Leonard Freißmuth, Matías Mattamala, Nived Chebrolu, Simon Schaefer, Stefan Leutenegger, Maurice Fallon |
IROS | 2 |
| 2024 | Evaluation and Deployment of LiDAR-based Place Recognition in Dense ForestsabstractMany LiDAR place recognition systems have been developed and tested specifically for urban driving scenarios. Their performance in natural environments such as forests and woodlands have been studied less closely. In this paper, we analyzed the capabilities of four different LiDAR place recognition systems, both handcrafted and learning-based methods, using LiDAR data collected with a handheld device and legged robot within dense forest environments. In particular, we focused on evaluating localization where there is significant translational and orientation difference between corresponding LiDAR scan pairs. This is particularly important for forest survey systems where the sensor or robot does not follow a defined road or path. Extending our analysis we then incorporated the best performing approach, Logg3dNet, into a full 6-DoF pose estimation system—introducing several verification layers for precise registration. We demonstrated the performance of our methods in three operational modes: online SLAM, offline multi-mission SLAM map merging, and relocalization into a prior map. We evaluated these modes using data captured in forests from three different countries, achieving 80 % of correct loop closures candidates with baseline distances up to 5 m, and 60 % up to 10 m. Video at: https://youtu.be/86l-oxjwmjY Haedam Oh, Nived Chebrolu, Matías Mattamala, Leonard Freißmuth, Maurice Fallon |
IROS | 3 |
| 2023 | MEM: Multi-Modal Elevation Mapping for Robotics and LearningabstractElevation maps are commonly used to represent the environment of mobile robots and are instrumental for locomotion and navigation tasks. However, pure geometric information is insufficient for many field applications that require appearance or semantic information, which limits their applicability to other platforms or domains. In this work, we extend a 2.5D robot-centric elevation mapping framework by fusing multi-modal information from multiple sources into a popular map representation. The framework allows inputting data contained in point clouds or images in a unified manner. To manage the different nature of the data, we also present a set of fusion algorithms that can be selected based on the information type and user requirements. Our system is designed to run on the GPU, making it real-time capable for various robotic and learning tasks. We demonstrate the capabilities of our framework by deploying it on multiple robots with varying sensor configurations and showcasing a range of applications that utilize multi-modal layers, including line detection, human detection, and colorization. Gian Erni, Jonas Frey, Takahiro Miki, Matías Mattamala, Marco Hutter 0001 |
IROS | 4 |
| 2021 | Learning Camera Performance Models for Active Multi-Camera Visual Teach and RepeatabstractIn dynamic and cramped industrial environments, achieving reliable Visual Teach and Repeat (VT&R) with a single-camera is challenging. In this work, we develop a robust method for non-synchronized multi-camera VT&R. Our contribution are expected Camera Performance Models (CPM) which evaluate the camera streams from the teach step to determine the most informative one for localization during the repeat step. By actively selecting the most suitable camera for localization, we are able to successfully complete missions when one of the cameras is occluded, faces into feature poor locations or if the environment has changed. Furthermore, we explore the specific challenges of achieving VT&R on a dynamic quadruped robot, ANYmal. The camera does not follow a linear path (due to the walking gait and holonomicity) such that precise path-following cannot be achieved. Our experiments feature forward and backward facing stereo cameras showing VT&R performance in cluttered indoor and outdoor scenarios. We compared the trajectories the robot executed during the repeat steps demonstrating typical tracking precision of less than 10 cm on average. With a view towards omni-directional localization, we show how the approach generalizes to four cameras in simulation. Matías Mattamala, Milad Ramezani, Marco Camurri, Maurice Fallon |
ICRA | 1 |
| 2020 | The Newer College Dataset: Handheld LiDAR, Inertial and Vision with Ground TruthabstractIn this paper, we present a large dataset with a variety of mobile mapping sensors collected using a handheld device carried at typical walking speeds for nearly 2.2 km around New College, Oxford as well as a series of supplementary datasets with much more aggressive motion and lighting contrast. The datasets include data from two commercially available devices - a stereoscopic-inertial camera and a multi-beam 3D LiDAR, which also provides inertial measurements. Additionally, we used a tripod-mounted survey grade LiDAR scanner to capture a detailed millimeter-accurate 3D map of the test location (containing ~290 million points). Using the map, we generated a 6 Degrees of Freedom (DoF) ground truth pose for each LiDAR scan (with approximately 3 cm accuracy) to enable better benchmarking of LiDAR and vision localisation, mapping and reconstruction systems. This ground truth is the particular novel contribution of this dataset and we believe that it will enable systematic evaluation which many similar datasets have lacked. The large dataset combines both built environments, open spaces and vegetated areas so as to test localisation and mapping systems such as vision-based navigation, visual and LiDAR SLAM, 3D LiDAR reconstruction and appearance-based place recognition, while the supplementary datasets contain very dynamic motions to introduce more challenges for visual-inertial odometry systems. The datasets are available at:ori.ox.ac.uk/datasets/newer-college-dataset. Milad Ramezani, Yiduo Wang 0001, Marco Camurri, David Wisth, Matías Mattamala, Maurice Fallon |
IROS | 5 |
| 2018 | Visual SLAM-Based Localization and Navigation for Service Robots: The Pepper CaseabstractWe propose a Visual-SLAM based localization and navigation system for service robots. Our system is built on top of the ORB-SLAM monocular system but extended by the inclusion of wheel odometry in the estimation procedures. As a case study, the proposed system is validated using the Pepper robot, whose short-range LIDARs and RGB-D camera do not allow the robot to self-localize in large environments. The localization system is tested in navigation tasks using Pepper in two different environments: a medium-size laboratory, and a large-size hall. Cristopher Gómez, Matías Mattamala, Tim Resink, Javier Ruiz-del-Solar |
RoboCup | 2 |
| 2017 | The NAO Backpack: An Open-Hardware Add-on for Fast Software Development with the NAO Robot
Matías Mattamala, Gonzalo Olave, Clayder Gonzalez-Cadenillas, Nicolás Hasbún, Javier Ruiz-del-Solar |
RoboCup | 1 |
| 2015 | A Dynamic and Efficient Active Vision System for Humanoid Soccer RobotsabstractThis paper presents an efficient active vision system which controls the head of a humanoid soccer robot. The system explicitly separates static information obtained offline from the map, and dynamic information from mobile objects, such as the ball and other players. Both types of information are mapped and handled in a simplified structure called action space , which assigns scores to each possible action of the robot’s head. Scores also consider the movement constraints of the robot’s head. Due to its simplicity and efficient information handling, the proposed active vision system is able to run in real-time in less than 1 ms. The performance of the system in a robot soccer environment is tested via simulation and real experiments. Matías Mattamala, Constanza Villegas, José Miguel Yáñez, Pablo Cano, Javier Ruiz-del-Solar |
RoboCup | 1 |