VLDB 2026 Research / reviewers in the wild / expert
Annette Stahl
dblp:28/5578
· DBLP profile ↗
22ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0002-8422-1091ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Near-Shore Mapping for Detection and Tracking of VesselsabstractFor an autonomous surface vessel (ASV) to dock, it must track other vessels close to the docking area. Kayaks present a particular challenge due to their proximity to the dock and relatively small size. Maritime target tracking has typically employed land masking to filter out land and the dock. However, imprecise land masking makes it difficult to track close-to-dock objects. Our approach uses Light Detection And Ranging (LiDAR) data and maps the docking area before tracking. The precise 3D measurements allow for precise map creation. However, the mapping could result in static, yet potentially moving, objects being mapped. We detect and filter out potentially moving objects from the LiDAR data by utilizing image data. The visual vessel detection and segmentation method is a neural network that is trained on our labeled data. Close-to-shore tracking improves with an accurate map and is demonstrated on a recently gathered real-world dataset. The dataset contains multiple sequences of a kayak and a day cruiser moving close to the dock, in a collision path with an autonomous ferry prototype. Nicholas Dalhaug, Annette Stahl, Rudolf Mester, Edmund Førland Brekke |
FUSION | 2 |
| 2025 | Stixel-Based Free Space Estimation for USVs Using Stereo Camera and LiDARabstractUnmanned surface vehicles (USVs) require robust situational awareness to navigate safely in complex maritime environments. A critical element of this is to identify the free navigable space around the USV. Free water regions can be derived from water segmentation in the image. However, these segmented regions must be transformed into a bird's eye view (BEV) representation to be utilized effectively in motion planning. This paper proposes a novel approach to estimate free navigable space in a BEV format by integrating a stereo camera and light detection and ranging (LiDAR). The proposed method uses water segmentation to delineate the water surface and represents the closest obstacles in the USV line of sight using vertical planar rectangles known as Stixels. The depth of these Stixels is derived from LiDAR data, ensuring precise positioning in space. The effectiveness of the approach is demonstrated through experiments conducted on real-world data collected from the milliAmpere 2 (MA2) autonomous ferry prototype in Trondheim, Norway. Qualitative evaluations focusing on accuracy and temporal consistency confirm its ability to reliably detect free navigable areas in complex maritime environments. Johannes Robert Skarø, Trym Anthonsen Nygård, Rudolf Mester, Annette Stahl, Edmund Førland Brekke |
FUSION | 4 |
| 2025 | Visual Lidar Recursive Online Tracker (ViLiROT) for Autonomous Surface VesselsabstractWe propose a multi-sensor fusion pipeline for multiple object tracking in autonomous surface vessels using lidar and camera data. Our approach follows the tracking-by-detection paradigm, leveraging the precision of lidar for accurate state estimation and camera data for robust association. The method addresses issues with false tracks from lidar returns by suppressing non-moving objects on the basis of optical flow. We compare the proposed pipeline against prior work, particularly in the use of lidar and stereo cameras as depth modalities, demonstrating its effectiveness in improving tracking performance. Henrik Hilmarsen, Nicholas Dalhaug, Trym Anthonsen Nygård, Edmund Førland Brekke, Annette Stahl, Rudolf Mester |
ICRA | 5 |
| 2024 | Combining Short and Wide Baseline Stereo Cameras for Improved Maritime Target TrackingabstractTarget tracking is essential for autonomous vehicles to avoid collisions. Using a stereo camera for the target tracking gives a dense representation of the targets, contrary to the the sparser data on typical radars and lidars. With a wider baseline stereo camera the depth measurements are more accurate, but the stereo matching challenge is greater, especially in the maritime domain with reflections on the water. Earlier classical methods of tracking using stereo cameras have often tracked targets by first doing water surface estimation and then finding objects perturbing the plane. The challenge is then to get a good estimate of the water surface plane while still having precise measurements to the targets. We propose both a short baseline method and a multi-baseline method for target detection. The multi-baseline method uses a short baseline stereo camera to find the water plane and uses a wider baseline stereo camera to get accurate target measurements. The targets are consistently being tracked when using data collected during the summer of 2023 from an autonomous ferry prototype compared to ground truth GNSS tracks. The short baseline method achieves minimal error for a day cruiser boat 40 m away using a camera baseline of only 12 cm. The multi-baseline method further improves the accuracy of boat measurements, especially for a far-away small kayak. Nicholas Dalhaug, Annette Stahl, Rudolf Mester, Edmund Førland Brekke |
FUSION | 2 |
| 2024 | FusedWSS: Water Surface Segmentation Fusing Machine Learning and Geometric CuesabstractNavigating unmanned surface vehicles (USVs) in urban waterways presents unique challenges due to irregular waterlines, obstacles, and reflections in the water. Determining the collision-free navigable area is crucial to enable safe USV operation. This paper introduces Fused Water Surface Segmentation (FusedWSS), a novel approach to water surface segmentation that aims to enhance navigation capabilities for USVs in complex harbor environments using a stereo camera. The method locates the water plane by performing plane fitting with outlier rejection and plane validation on the reconstructed 3D point cloud. From the plane parameters, the virtual horizon line is inferred and used for point cloud and image cropping. The water surface mask and virtual horizon line are fused with a deep learningbased semantic segmentation method to produce accurate and reliable water masks for each image frame. Additional refinement of the water mask is performed using detected obstacle masks. Validation was carried out using data from the MilliAmpere 2 autonomous ferry prototype in Trondheim, Norway, and a publicly available maritime dataset, demonstrating the efficacy of the methods. Jon Torgeir Grini, Rudolf Mester, Trym Anthonsen Nygård, Nicholas Dalhaug, Edmund Førland Brekke, Annette Stahl |
FUSION | 6 |
| 2024 | Maritime Tracking-By-Detection with Object Mask Depth Retrieval Through Stereo Vision and LidarabstractThe momentum towards autonomous technology is building up in the maritime domain, as the automotive industry has made big steps towards autonomous driving. The automotive industry has increasingly utilized visual methods for multi-object tracking (MOT), with the help of accessible benchmarking datasets such as KITTI. This paper presents a tracking pipeline that tracks in the world frame by using elements of a well-established visual tracking method that tracks objects in the image frame. The pipeline fuses 3D information from lidar or stereo vision with object masks from a deep learning-based ship detector. To handle occlusions, we implemented a track manager that predicts lost objects’ movement until they reappear. Also, we provide a comparison between using lidar and stereo as the depth modality in the tracking pipeline. Results from a real-world experiment indicate that camera-lidar fusion gives consistently precise estimates, while the precision with stereo depends on the range and the type of vessel tracked. Henrik Hilmarsen, Nicholas Dalhaug, Trym Anthonsen Nygård, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 6 |
| 2024 | A General Low-Parameter 3D Ship Hull Extent Model for Object TrackingabstractIn autonomous vehicle systems, it is paramount to detect other objects in the vicinity and track their movement. Extended Object Tracking (EOT) provides a convenient framework for tracking objects using high-resolution sensor data by defining models for the object’s spatial dimensions (a.k.a. extent). In maritime applications, the objects of interest are mainly other maritime vessels, and these vary greatly in shape and size. This diversity proves to be a challenge for defining general extent models that both give accurate representations for most vessels and that do not depend on a large number of parameters. In this paper, a general three-dimensional low-parameter ship hull model designed for EOT is presented. The presented extent model is constructed by intertwining a polynomial representation along the vertical direction with a frequency representation along the horizontal plane. However, to reduce the dimension of the parameter space without compromising its accuracy, the horizontal frequency representation is modified by performing a Principal Component Analysis (PCA). In particular, this extent representation does not require an underlying discretization grid, which makes the model scalable and therefore well-suited for modeling objects that vary greatly in size. Michael Ernesto López, Kjetil Vasstein, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 5 |
| 2024 | RGB-D Mapping and Tracking in a Plenoxel Radiance FieldabstractThe widespread adoption of Neural Radiance Fields (NeRFs) have ensured significant advances in the domain of novel view synthesis in recent years. These models capture a volumetric radiance field of a scene, creating highly convincing, dense, photorealistic models through the use of simple, differentiable rendering equations. Despite their popularity, these algorithms suffer from severe ambiguities in visual data inherent to the RGB sensor, which means that although images generated with view synthesis can visually appear very believable, the underlying 3D model will often be wrong. This considerably limits the usefulness of these models in practical applications like Robotics and Extended Reality (XR), where an accurate dense 3D reconstruction otherwise would be of significant value. In this paper, we present the vital differences between view synthesis models and 3D reconstruction models. We also comment on why a depth sensor is essential for modeling accurate geometry in general outward-facing scenes using the current paradigm of novel view synthesis methods. Focusing on the structure-from-motion task, we practically demonstrate this need by extending the Plenoxel radiance field model: Presenting an analytical differential approach for dense mapping and tracking with radiance fields based on RGB-D data without a neural network. Our method achieves state-of-the-art results in both mapping and tracking tasks, while also being faster than competing neural network-based approaches. The code is available at: https://github.com/ysus33/RGB-D_Plenoxel_Mapping_Tracking.git. Andreas Langeland Teigen, Yeonsoo Park, Annette Stahl, Rudolf Mester |
WACV | 3 |
| 2024 | Robust Hole-Detection in Triangular Meshes Irrespective of the Presence of Singular VerticesabstractIn this work, we present a boundary and hole detection approach that traverses all the boundaries of an edge-manifold triangular mesh, irrespectively of the presence of singular vertices, and subsequently determines and labels all holes of the mesh. The proposed automated hole-detection method is valuable to the computer-aided design (CAD) community as all boundary-edges within the mesh are utilized and for each boundary-edge the algorithm guarantees both the existence and the uniqueness of the boundary associated to it. As existing hole-detection approaches assume that singular vertices are absent or may require mesh modification, these methods are ill-equipped to detect boundaries/holes in real-world meshes that contain singular vertices. We demonstrate the method in an underwater autonomous robotic application, exploiting surface reconstruction methods based on point cloud data. In such a scenario the determined holes can be interpreted as information gaps, enabling timely corrective action during the data acquisition. However, the scope of our method is not confined to these two sectors alone; it is versatile enough to be applied on any edge-manifold triangle mesh. An evaluation of the method is performed on both synthetic and real-world data (including a triangle mesh from a point cloud obtained by a multibeam sonar). The source code of our reference implementation is available: https://github.com/Mauhing/hole-detection-on-triangle-mesh. Mauhing Yip, Annette Stahl, Christian Schellewald |
Comput. Aided Des. | 2 |
| 2023 | Maritime radar odometry inspired by visual odometryabstractFuture autonomous ships will need several redundant positioning systems to navigate reliably. Global Navigation Satellite Systems are highly accurate but they are susceptible to disruptions and intentional jamming. Maritime radars have long range and are robust against bad weather and darkness, but the use for ownship motion estimation has received relatively little attention in the research field. In this work, we present a radar odometry estimation method inspired by advances in visual odometry and simultaneous localization and mapping. The method works on raw radar data in a coastal environment and combines the Kanade-Lucas-Tomashi tracker with a factor graph back-end. We test it on data from a large ship with a maritime radar with a range of 19 km. We find that it is robust with only a small drift and no erroneous jumps in the estimate. Henrik D. Flemmen, Rudolf Mester, Annette Stahl, Torleiv H. Bryne, Edmund Førland Brekke |
FUSION | 3 |
| 2023 | Multiscan Shape Estimation for Extended Object TrackingabstractExtended Object Tracking (EOT) is a advantageous technique for achieving situational awareness in autonomous vehicle systems. The EOT problem is to both estimate the movement and spatial dimensions of an object using high-resolution measurements. In the case of laser measurements or other types of measurements that correspond to points on the object’s boundary, the true measurement model of the EOT problem is based on an implicit equation for the measurement coordinates. This intrinsic implicity is often not addressed directly in several EOT models found in the literature. In this paper, the EOT problem is reformulated as a least square minimization problem without compromising the original implicit measurement model by introducing an extra variable for each measurement. In addition, this new least squares formulation allows considering measurements and state variables for a whole time window, and not just a single time step. An EOT algorithm based on solving the derived least squares minimization problem is proposed and tested with simulated scenarios. Michael Ernesto López, Edmund Førland Brekke, Rudolf Mester, Annette Stahl |
FUSION | 4 |
| 2021 | MOG: a background extraction approach for data augmentation of time-series images in deep learning segmentationabstractImage segmentation is one of the key components in systems performing computer vision recognition tasks. Various algorithms for image segmentation have been developed in the literature. Among them, more recently, deep learning algorithms have been remarkably successful in performing this task. A downside with deep neural networks for segmentation is that they require a large amount of labeled dataset for training. This prerequisite is one of the main reasons that led researchers to adopt data augmentation approaches in order to minimize manual labeling efforts while maintaining highly accurate results. This paper uses classical non-deep learning methods for background extraction to increase the size of the dataset used to train deep learning attention segmentation algorithms when images are presented as time-series to the model. The method presented adopts the Gaussian mixture-based (MOG2) foreground-background segmentation followed by dilation and erosion to create masks necessary to train the deep learning models. It is applied in the context of planktonic images captured in situ as time series. Various evaluation metrics and visual inspection are used to compare the performance of the deep learning algorithms. Experimental results show higher accuracy achieved by the deep learning algorithms for time-series image attention segmentation when the proposed data augmentation methodology is utilized to increase the training dataset. Jonas Nagell Borgersen, Aya Saad, Annette Stahl |
ICMV | 3 |
| 2021 | Robust deep unsupervised learning framework to discover unseen plankton speciesabstractDeep convolutional neural networks have proven effective in computer vision, especially in the task of image classification Nevertheless, the success is limited to supervised learning approaches, requiring extensive amounts of labeled training data that impose time-consuming manual efforts. Unsupervised deep learning methods were introduced to overcome this challenge. The gap, however, towards achieving comparable classification accuracy to supervised learning is still significant. This paper presents a deep learning framework for images of planktonic organisms with no ground truth or manually labeled data. This work combines feature extraction methods using state-of-the-art unsupervised training schemes with clustering algorithms to minimize the labeling effort while improving the classification process based on essential features learned by the deep learning model. The models utilized in the framework are tested over existing planktonic data sets. Empirical results show that unsupervised approaches that cluster the data based on the deep learning model’s feature space representations improve the classification task and can identify classes that have not been seen during the learning process. Eivind Salvesen, Aya Saad, Annette Stahl |
ICMV | 3 |
| 2021 | Progressive integration of visibility constraints for implicit functionsabstractProcedurally-defined implicit functions, such as CSG trees and recent neural shape representations, offer compelling benefits for modeling scenes and objects, including infinite resolution, differentiability and trivial deformation, at a low memory footprint. The common approach to fit such models to measurements is to solve an optimization problem involving the function evaluated at points in space. However, the computational cost of evaluating the function makes it challenging to use visibility information from range sensors and 3D reconstruction systems. We propose a method that uses visibility information, where the number of function evaluations required at each iteration is proportional to the scene area. Our method builds on recent results for bounded Euclidean distance functions by introducing a coarse-to-fine mechanism to avoid the requirement for correct bounds. This makes our method applicable to a greater variety of implicit modeling techniques, for which deriving the Euclidean distance function or appropriate bounds is difficult. Annette Stahl |
ICMV | 1 |
| 2021 | Hole detection in aquaculture net cages from video footageabstractFrequent inspection of salmon cage integrity is essential to early detect and prevent the possible escape of farmed salmon—minimizing the risk of any negative impact for the remaining wild stock of salmon. Current state-of-the-art computer vision-based approaches can detect net irregularities under “optimal” net and illumination conditions but might fail under real-world conditions. In this paper, we present a novel modularized processing framework based on advanced computer vision and machine learning approaches to effectively detect potential net damages in video recordings from cleaner robots traversing the net cages. The framework includes a deep learning-based approach to segmenting interpretable net structure from background, transfer learning facilitated classification of potential holes from irrelevance, and computer vision-based modules for irregularity detection, filtering, and tracking. Filtering and classification are vital steps to ensure that temporally consistent holes within net structure are reported—and irrelevant objects such as by-passing fish are ignored. We evaluate our approach on representative real-world videos from real cleaning operations and show that the approach can cope with the difficult lighting conditions that are typical for aquaculture environments. Annette Stahl |
ICMV | 1 |
| 2020 | Feature-Based Laser Odometry for Autonomous Surface Vehicles utilizing the Point Cloud LibraryabstractThis paper proposes a pipeline for feature-based laser odometry for autonomous surface vehicles (ASVs) operating in urban environments. In particular, we investigate the suitability of several keypoint extractors and keypoint descriptors available through the Point Cloud Library (PCL). The complete odometry system, using the different extractors and descriptors, is implemented using the iSAM2 framework, and validated on real lidar data recorded onboard the autonomous ferry prototype MilliAmpere. The results demonstrate that accuracy similar to a standard Global Navigation Satellite System (GNSS) receiver can be achieved after 10 minutes even without loop closure. This study can be used as a starting point for future research on Simultaneous Localization And Mapping (SLAM) for ASVs. Even Skjellaug, Edmund Førland Brekke, Annette Stahl |
FUSION | 3 |
| 2020 | An instance segmentation framework for in-situ plankton taxa assessmentabstractIn this paper, we propose a deep learning instance segmentation framework for particle extraction of microscopic images that aims at calculating planktonic species distribution and concentration in-situ. The framework comprises three essential functional tasks on in-situ time-series images collected from an autonomous underwater vehicle: 1) manual labeling of the captured images, 2) object localization, segmentation, and identification, and 3) class distribution and planktonic organisms concentration calculation. Our proposed framework is based on the mask R-CNN architecture provided by the Detectron2 library developed by Facebook Artificial Intelligence Research (FAIR) for instance segmentation. Due to its modular design, we compare the performance of different networks by alternating the backbone sub-network in order to choose the most suitable architecture for the task of instance and semantic segmentation. We compile a custom annotated dataset from planktonic time-series images and train the different models over this dataset to perform the instance semantic segmentation. Evaluation results of the proposed framework, utilizing the best performing deep learning architecture along with the new annotated dataset, show better performance in terms of speed and accuracy of both in-situ segmentation and classification compared to traditional segmentation methods. In addition, we observe a significant improvement in the object classification quality when we train the model over our newly annotated dataset instead of training it over the dataset generated from the traditional methods. The inferred data from our novel instance segmentation framework, which provides the particle class distribution and concentration, can then be used to assist in constructing a dynamic probability density map of planktonic communities dispersion and abundance. Aya Saad, Annette Stahl |
ICMV | 2 |
| 2020 | Image Segmentation of Corrosion Damages in Industrial InspectionsabstractIn this paper we provide insights and methods for using image segmentation for the purposes of automatic corrosion damage detection. Automatic image analysis is needed in order to process all data retrieved from drone-driven industrial inspections. To this end we provide three main contributions. First, 608 images with corrosion damages are instance-wise annotated with binary segmentation masks to construct a dataset. Second, a novel, two-stage data augmentation scheme is developed and empirically shown to significantly reduce overfitting. Finally, Mask R-CNN and PSPNet are evaluated on the corrosion dataset using this and other data augmentation methods. With 77.5% and 73.2% mean intersection over union (mIoU) for Mask R-CNN and PSPNet, respectively, the results are very promising. It is concluded that image segmentation can aid automating industrial inspections of steel constructions in the future, and that instance segmentation is likely more useful than semantic segmentation due to its applications to a wider range of use-cases. Simen Keiland Fondevik, Annette Stahl, Aksel Andreas Transeth, Ole Ø. Knudsen |
ICTAI | 2 |
| 2019 | Classification of corrosion and coating damages on bridge constructions from images using convolutional neural networksabstractIn this paper, we present a comparison of performance for different convolutional neural networks (CNN) for automatic classification of corrosion and coating damages on bridge constructions from images. Image recordings were taken during inspections. Through manual categorization and data augmentation, a total of 9300 images were collected and divided into five classes. Four different CNNs were trained using transfer learning in MATLAB. We have evaluated test performance through the metrics recall, precision, accuracy and F1 score. Test performance was also evaluated on damage detection accuracy, meaning how well the networks detect images that contain a damage. The convolutional neural network trained using VGG-16 had the overall best performance results, with average recall, precision, accuracy and F1 score being 95.45%, 95.61%, 97.74% and 95.53%, respectively. In the category of overall damage detection AlexNet performed best with 99.14% accuracy. The obtained results are promising, and make it possible to conclude that CNNs have a great potential in bridge inspections for automatic analysis of corrosion and coating damages. Egil Holm, Aksel Andreas Transeth, Ole Ø. Knudsen, Annette Stahl |
ICMV | 4 |
| 2018 | Teaching a Robot to Grasp Real Fish by Imitation Learning from a Human Supervisor in Virtual RealityabstractWe teach a real robot to grasp real fish, by training a virtual robot exclusively in virtual reality. Our approach implements robot imitation learning from a human supervisor in virtual reality. A deep 3D convolutional neural network computes grasps from a 3D occupancy grid obtained from depth imaging at multiple viewpoints. In virtual reality, a human supervisor can easily and intuitively demonstrate examples of how to grasp an object, such as a fish. From a few dozen of these demonstrations, we use domain randomization to generate a large synthetic training data set consisting of 100 000 example grasps of fish. Using this data set for training purposes, the network is able to guide a real robot and gripper to grasp real fish with good success rates. The newly proposed domain randomization approach constitutes the first step in how to efficiently perform robot imitation learning from a human supervisor in virtual reality in a way that transfers well to the real world. Jonatan S. Dyrstad, Elling Ruud Øye, Annette Stahl, John Reidar Mathiassen |
IROS | 3 |
| 2017 | Continuous Signed Distance Functions for 3D VisionabstractWe explore the use of continuous signed distance functions as an object representation for 3D vision. Popularized in procedural computer graphics, this representation defines 3D objects as geometric primitives combined with constructive solid geometry and transformed by nonlinear deformations, scaling, rotation or translation. Unlike its discretized counterpart, that has become important in dense 3D reconstruction, the continuous distance function is not stored as a sampled volume, but as a closed mathematical expression. We argue that this representation can have several benefits for 3D vision, such as being able to describe many classes of indoor and outdoor objects at the order of hundreds of bytes per class, getting parametrized shape variations for free. As a distance function, the representation also has useful computational aspects by defining, at each point in space, the direction and distance to the nearest surface, and whether a point is inside or outside the surface. Simen Haugo, Annette Stahl, Edmund Førland Brekke |
3DV | 2 |
| 2010 | Image Motion Estimation using Optimal Flow Control
Annette Stahl, Ole Morten Aamo |
ICINCO (3) | 1 |