VLDB 2026 Research / reviewers in the wild / expert
Simone Frintrop
dblp:99/192
· DBLP profile ↗
54ranked-venue papers
13as first author
18since 2021 · last 2025
0000-0002-9475-3593ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 37 · 11 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 13 since 2021Systems, architecture and hardware · 17 · 6 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Walk the Lines 2: Contour Tracking for Detailed Segmentation of Infrared Ships and Other Objects
André Peter Kelm, Max Braeschke, Emre Gülsoylu, Simone Frintrop |
CAIP (1) | 4 |
| 2025 | Contextloss: Context Information for Topology-Preserving SegmentationabstractIn image segmentation, preserving the topology of segmented structures like vessels, membranes, or roads is crucial. For instance, topological errors on road networks can significantly impact navigation. Recently proposed solutions are loss functions based on critical pixel masks that consider the whole skeleton of the segmented structures in the critical pixel mask. We propose the novel loss function ContextLoss (CLoss) that improves topological correctness by considering topological errors with their whole context in the critical pixel mask. The additional context improves the network focus on the topological errors. Further, we propose two intuitive metrics to verify improved connectivity due to a closing of missed connections. We benchmark our proposed CLoss on three public datasets (2D & 3D) and our own 3D nano-imaging dataset of bone cement lines. Training with our proposed CLoss increases performance on topology-aware metrics and repairs up to 44 % more missed connections than other state-of-the-art methods. We make the code publicly available12. Benedict Schacht, Imke Greving, Simone Frintrop, Berit Zeller-Plumhoff, Christian Wilms |
ICIP | 3 |
| 2024 | SOS: Segment Object System for Open-World Instance Segmentation with Object Priors
Christian Wilms, Tim Rolff, Maris Hillemann, Robert Johanson, Simone Frintrop |
ECCV (27) | 5 |
| 2024 | Dynamic Inference and Top-down Attention in a Hierarchical Classification Network
André Peter Kelm, Niels Hannemann, Bruno Heberle, Lucas Schmidt, Tim Rolff, Christian Wilms, Ehsan Yaghoubi, Simone Frintrop |
ICPR (8) | 8 |
| 2024 | S3AD: Semi-supervised Small Apple Detection in Orchard EnvironmentsabstractCrop detection is integral for precision agriculture applications such as automated yield estimation or fruit picking. However, crop detection, e.g., apple detection in orchard environments remains challenging due to a lack of large-scale datasets and the small relative size of the crops in the image. In this work, we address these challenges by reformulating the apple detection task in a semi-supervised manner. To this end, we provide the large, high-resolution dataset MAD1comprising 105 labeled images with 14,667 annotated apple instances and 4,440 unlabeled images. Utilizing this dataset, we also propose a novel Semi-Supervised Small Apple Detection system S3AD based on contextual attention and selective tiling to improve the challenging detection of small apples, while limiting the computational overhead. We conduct an extensive evaluation on MAD and the MSU dataset, showing that S3AD substantially outperforms strong fully-supervised baselines, including several small object detection systems, by up to 14.9%. Additionally, we exploit the detailed annotations of our dataset w.r.t. apple properties to analyze the influence of relative size or level of occlusion on the results of various systems, quantifying current challenges. Robert Johanson, Christian Wilms, Ole Johannsen, Simone Frintrop |
WACV | 4 |
| 2023 | A Deep Learning Architecture for Egocentric Time-to-Saccade Prediction using Weibull Mixture-Models and Historic PriorsabstractReal-time detection of saccades is of major interest for many applications in human-computer interaction and mixed reality. However, due to relatively low update rates and high latencies of current commercially available eye trackers, gaze events are typically detected after they occur with some delay. This limits interaction scenarios such as intent-based gaze interaction, redirected walking, or gaze forecasting. Tim Rolff, Susanne Schmidt 0001, Frank Steinicke, Simone Frintrop |
ETRA | 4 |
| 2023 | Hands in Focus: Sign Language Recognition Via Top-Down AttentionabstractIn this paper, we propose a novel Sign Language Recognition (SLR) model that leverages the task-specific knowledge to incorporate Top-Down (TD) attention to focus the processing of the network on the most relevant parts of the input video sequence. For SLR, this includes information about the hands’ shape, orientation and positions, and motion trajectory. Our model consists of three streams that process RGB, optical flow and TD attention data. For the TD attention, we generate pixel-precise attention maps focusing on both hands, thereby retaining valuable hand information, while eliminating distracting background information. Our proposed method outperforms state-of-the-art on a challenging large-scale dataset by over 2%, and achieves strong results with a much simpler architecture compared to other systems on the newly released AUTSL dataset [1]. Noha A. Sarhan, Christian Wilms, Vanessa Closius, Ulf Brefeld, Simone Frintrop |
ICIP | 5 |
| 2023 | Small, but Important: Traffic Light Proposals for Detecting Small Traffic Lights and Beyond
Tom Sanitz, Christian Wilms, Simone Frintrop |
ICVS | 3 |
| 2023 | VRS-NeRF: Accelerating Neural Radiance Field Rendering with Variable Rate ShadingabstractRecent advancements in Neural Radiance Fields (NeRF) provide enormous potential for a wide range of Mixed Reality (MR) applications. However, the applicability of NeRF to real-time MR systems is still largely limited by the rendering performance of NeRF. In this paper, we present a novel approach for Variable Rate Shading for Neural Radiance Fields (VRS-NeRF). In contrast to previous techniques, our approach does not require training multiple neural networks or re-training of already existing ones, but instead utilizes the raytracing properties of NeRF. This is achieved by merging rays depending on a variable shading rate, which reduces the overall number of queries to the neural network. We demonstrate the generalizability of our approach by implementing three alternative functions for the determination of the shading rate. The first method uses the gaze of users to effectively implement a foveated rendering technique in NeRF. For the other two techniques, we utilize shading rates based on edges and saliency. Based on a psychophysical experiment and multiple image-based metrics, we suggest a set of parameters for each technique, yielding an optimal tradeoff between rendering performance gain and perceived visual quality. Tim Rolff, Susanne Schmidt 0001, Ke Li 0025, Frank Steinicke, Simone Frintrop |
ISMAR | 5 |
| 2022 | When do Saccades begin? Prediction of Saccades as a Time-to-Event ProblemabstractWe present a novel view on gaze event classification by redefining it as a time-to-event problem. In contrast to previous models, which consider the classification as discrete events, our redefinition allows for estimating the remaining time until the next saccade event. Therefore, we provide a feature analysis and an initial solution for compensating the latency of wearable eye-trackers build in today’s head-mounted displays. Similar to previous classifiers, we utilize oculomotor features such as velocity, acceleration, and event durations. In total, we analyze 104 extracted features of three datasets and apply different regression methods. We identify optimal window sizes for each feature and extract the importance of all extracted windows using recursive feature elimination. Afterwards, we evaluate the performance of all regressors using earlier selected features. We show that our selected regressors can predict the time-to-event better than the baseline, indicating the potential usage of time-to-event prediction of saccades. Tim Rolff, Frank Steinicke, Simone Frintrop |
ETRA | 3 |
| 2022 | HD Ground - A Database for Ground Texture Based LocalizationabstractWe present the HD Ground Database, a comprehensive database for ground texture based localization. It contains sequences of a variety of textures, obtained using a downward facing camera. In contrast to existing databases of ground images, the HD Ground Database is larger, has a greater variety of textures, and has a higher image resolution with less motion blur. Also, our database enables the first systematic study of how natural changes of the ground that occur over time affect localization performance, and it allows to examine a teach-and-repeat navigation scenario. We use the HD Ground Database to evaluate four state-of-the-art localization approaches for global localization, localization with the approximate pose being known, and relative localization. Jan Fabian Schmid, Stephan F. Simon, Raaghav Radhakrishnan, Simone Frintrop, Rudolf Mester |
ICRA | 4 |
| 2022 | The MSR-Video to Text dataset with clean annotations
Haoran Chen 0011, Jianmin Li 0001, Simone Frintrop, Xiaolin Hu 0001 |
Comput. Vis. Image Underst. | 3 |
| 2021 | Sign, Attend and Tell: Spatial Attention for Sign Language RecognitionabstractSign Language Recognition (SLR) has witnessed a boost in recent years, particularly with the surge of deep learning techniques. However, most existing methods do not exploit the concept of attention mechanisms, despite their success in several computer vision tasks. In this paper, we propose a novel method for isolated SLR which utilizes spatial attention to focus the processing on the informative, discriminating parts of the input. This is particularly important for SLR, since the RGB image contains several distracting information such as background and signer's clothes, which are irrelevant for the task. We investigate three ways for incorporating spatial attention: a) pre-focused attention, which uses optical-flow-based motion as a prior b) learned attention, where the network learns where to focus during training, and c) hybrid attention, which combines both approaches by initializing the attention layer in the learned attention with the motion-based attention masks used in the pre-focused attention. We show, first, that all three approaches outperform state-of-the-art methods on one of the largest isolated SLR datasets, validating the effectiveness of attention mechanisms on the SLR task, and second, that the best performing approach is the hybrid attention, combining both ideas. Noha A. Sarhan, Simone Frintrop |
FG | 2 |
| 2021 | CloudAAE: Learning 6D Object Pose Regression with On-line Data Synthesis on Point CloudsabstractIt is often desired to train 6D pose estimation systems on synthetic data because manual annotation is expensive. However, due to the large domain gap between the synthetic and real images, synthesizing color images is expensive. In contrast, this domain gap is considerably smaller and easier to fill for depth information. In this work, we present a system that regresses 6D object pose from depth information represented by point clouds, and a lightweight data synthesis pipeline that creates synthetic point cloud segments for training. We use an augmented autoencoder (AAE) for learning a latent code that encodes 6D object pose information for pose regression. The data synthesis pipeline only requires texture-less 3D object models and desired viewpoints, and it is cheap in terms of both time and hardware storage. Our data synthesis process is up to three orders of magnitude faster than commonly applied approaches that render RGB image data. We show the effectiveness of our system on the LineMOD, LineMOD Occlusion, and YCB Video datasets. The implementation of our system is available at: https://github.com/GeeeG/CloudAAE. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 5 |
| 2021 | Object Localization with Attribute Preference Based on Top-Down Attention
Soubarna Banik, Mikko Lauri, Alois C. Knoll, Simone Frintrop |
ICVS | 4 |
| 2021 | See the Silence: Improving Visual-Only Voice Activity Detection by Optical Flow and RGB Fusion
Danu Caus, Guillaume Carbajal, Timo Gerkmann, Simone Frintrop |
ICVS | 4 |
| 2021 | Guest Editorial: Special Issue: Computer Vision and Pattern Recognition (DAGM GCPR 2019)
Simone Frintrop, Gernot A. Fink, Xiaoyi Jiang 0001 |
Int. J. Comput. Vis. | 1 |
| 2021 | DeepFH segmentations for superpixel-based object proposal refinement
Christian Wilms, Simone Frintrop |
Image Vis. Comput. | 2 |
| 2020 | Transfer Learning For Videos: From Action Recognition To Sign Language RecognitionabstractIn this paper, we propose using Inflated 3D (I3D) Convolutional Neural Networks for large-scale signer-independent sign language recognition (SLR). Unlike other recent methods, our method relies only on RGB video data and does not require other modalities such as depth. This is beneficial for many applications in which depth data is not available. We show that transferring spatiotemporal features from a large-scale action recognition dataset is highly valuable to the training for SLR. Based on an architecture for action recognition [1], we use two-stream I3D ConvNets operating on RGB and optical flow images. Our method is evaluated on the ChaLearn249 Isolated Gesture Recognition dataset and clearly outperforms other state-of-the-art RGB-based methods. Noha A. Sarhan, Simone Frintrop |
ICIP | 2 |
| 2020 | Improving mix-and-separate training in audio-visual sound source separation with an object priorabstractThe performance of an audio-visual sound source separation system is determined by its ability to separate audio sources given the images of the sources and the audio mixture. The goal of this study is to investigate the ability to learn the mapping between the sounds and the images of instruments in the self-supervisied mix-and-seperate training paradigm used by state-of-the-art audio-visual sound source separation methods. Theoretical and empirical analyses illustrate that the self-supervised mix-and-separate training does not automatically learn the 1-to-1 correspondence between visual and audio signals, leading to low audio-visual object classification accuracy. Based on this analysis, a weakly-supervised method called Object-Prior is proposed and evaluated on two audio-visual datasets. The experimental results show that the Object-Prior method outperforms state-of-the-art baselines in the audio-visual sound source separation task. It is also more robust against asynchronized data, where the frame and the audio do not come from the same video, and recognizes musical instruments based on their sound with higher accuracy. This indicates that learning the 1-to-1 correspondence between visual and audio features of an instrument improves the effectiveness of audio-visual sound source separation. Julius Richter, Mikko Lauri, Timo Gerkmann, Simone Frintrop |
ICPR | 5 |
| 2020 | Superpixel-based Refinement for Object Proposal GenerationabstractPrecise segmentation of objects is an important problem in tasks like class-agnostic object proposal generation or instance segmentation. Deep learning-based systems usually generate segmentations of objects based on coarse feature maps, due to the inherent downsampling in CNNs. This leads to segmentation boundaries not adhering well to the object boundaries in the image. To tackle this problem, we introduce a new superpixel-based refinement approach1on top of the state-of-the-art object proposal system AttentionMask. The refinement utilizes superpixel pooling for feature extraction and a novel superpixel classifier to determine if a high precision superpixel belongs to an object or not. Our experiments show an improvement of up to 26.0% in terms of average recall compared to original AttentionMask. Furthermore, qualitative and quantitative analyses of the segmentations reveal significant improvements in terms of boundary adherence for the proposed refinement compared to various deep learning-based state-of-theart object proposal generation systems. Christian Wilms, Simone Frintrop |
ICPR | 2 |
| 2020 | Which Airline is This? Airline Logo Detection in Real-World Weather ConditionsabstractThe detection of logos in images, for instance, logos of airlines on airplane tails, is a difficult task in real-world weather conditions. Most systems used for logo detection are very good at detecting logos in clean images. However, they exhibit problems when images are degraded by effects of adverse weather conditions as they frequently occur in real-world scenarios. For investigating this problem on airline logo detection as a subproblem of logo detection, we first present a new dataset for airline logo detection on airplane tails containing a test split with images degraded by adverse weather effects. Second, to handle the detection of airline logos effectively, a new two-stage airline logo detection system based on a state-of-the-art object proposal generation system and a specifically tailored classifier is proposed. Finally, improving the results on images degraded by adverse weather effects, we introduce a learning-free application-agnostic data augmentation strategy simulating effects like rain and fog. The results show the superior performance of our airline logo detection system compared to state-of-the-art. Furthermore, applying our data augmentation approach to a variety of systems, reduces the significant drop in performance on degraded images. Christian Wilms, Rafael Heid, Mohammad Araf Sadeghi, Andreas Ribbrock, Simone Frintrop |
ICPR | 5 |
| 2020 | 6D Object Pose Regression via Supervised Learning on Point CloudsabstractThis paper addresses the task of estimating the 6 degrees of freedom pose of a known 3D object from depth information represented by a point cloud. Deep features learned by convolutional neural networks from color information have been the dominant features to be used for inferring object poses, while depth information receives much less attention. However, depth information contains rich geometric information of the object shape, which is important for inferring the object pose. We use depth information represented by point clouds as the input to both deep networks and geometry-based pose refinement and use separate networks for rotation and translation regression. We argue that the axis-angle representation is a suitable rotation representation for deep learning, and use a geodesic loss function for rotation regression. Ablation studies show that these design choices outperform alternatives such as the quaternion representation and L2 loss, or regressing translation and rotation with the same network. Our simple yet effective approach clearly outperforms state-of-the-art methods on the YCB-video dataset. Mikko Lauri, Xiaolin Hu 0001, Jianwei Zhang 0001, Simone Frintrop |
ICRA | 6 |
| 2019 | Explore, Approach, and Terminate: Evaluating Subtasks in Active Visual Object Search Based on Deep Reinforcement LearningabstractSearching for objects and distinguishing task-relevant objects from others is a key requirement for service robots. We propose a reinforcement learning solution to the active visual object search problem. Our method successfully learns to explore the environment, to approach the target object, and to decide when to terminate the search as the target object has been found. We demonstrate the efficiency of our solution on a dataset of real-world images collected by a robot. Our approach outperforms state-space planning or other baseline search strategies, reaching a higher success rate in a shorter time. We also study individual subtasks of active visual object search. Although strong baselines exist for the subtasks, our RL solution outperforms them in the overall search task. Jan Fabian Schmid, Mikko Lauri, Simone Frintrop |
IROS | 3 |
| 2019 | Distance Dependent Maximum Margin Dirichlet Process Mixture
Mikko Lauri, Simone Frintrop |
PRICAI (2) | 3 |
| 2018 | AttentionMask: Attentive, Efficient Object Proposal Generation Focusing on Small Objects
Christian Wilms, Simone Frintrop |
ACCV (2) | 2 |
| 2017 | Multi-robot active information gathering with periodic communicationabstractA team of robots sharing a common goal can benefit from coordination of the activities of team members, helping the team to reach the goal more reliably or quickly. We address the problem of coordinating the actions of a team of robots with periodic communication capability executing an information gathering task. We cast the problem as a multi-agent optimal decision-making problem with an information theoretic objective function. We show that appropriate techniques for solving decentralized partially observable Markov decision processes (Dec-POMDPs) are applicable in such information gathering problems. We quantify the usefulness of coordinated information gathering through simulation studies, and demonstrate the feasibility of the method in a real-world target tracking domain. Mikko Lauri, Eero Heinänen, Simone Frintrop |
ICRA | 3 |
| 2017 | Saliency-guided adaptive seeding for supervoxel segmentationabstractWe propose a new saliency-guided method for generating supervoxels in 3D space. Rather than using an evenly distributed spatial seeding procedure, our method uses visual saliency to guide the process of supervoxel generation. This results in densely distributed, small, and precise supervoxels in salient regions which often contain objects, and larger supervoxels in less salient regions that often correspond to background. Our approach largely improves the quality of the resulting supervoxel segmentation in terms of boundary recall and under-segmentation error on publicly available benchmarks. Mikko Lauri, Jianwei Zhang 0001, Simone Frintrop |
IROS | 4 |
| 2016 | Semantic segmentation priors for object discoveryabstractReliable object discovery in realistic indoor scenes is a necessity for many computer vision and service robot applications. In these scenes, semantic segmentation methods have made huge advances in recent years. Such methods can provide useful prior information for object discovery by removing false positives and by delineating object boundaries. We propose a novel method that combines bottom-up object discovery and semantic priors for producing generic object candidates in RGB-D images. We use a deep learning method for semantic segmentation to classify colour and depth superpixels into meaningful categories. Separately for each category, we use saliency to estimate the location and scale of objects, and superpixels to find their precise boundaries. Finally, object candidates of all categories are combined and ranked. We evaluate our approach on the NYU Depth V2 dataset and show that we outperform other state-of-the-art object discovery methods in terms of recall. Germán Martín García, Farzad Husain, Hannes Schulz, Simone Frintrop, Carme Torras, Sven Behnke |
ICPR | 4 |
| 2015 | Traditional saliency reloaded: A good old model in new shapeabstractIn this paper, we show that the seminal, biologically-inspired saliency model by Itti et al. [21] is still competitive with current state-of-the-art methods for salient object segmentation if some important adaptions are made. We show which changes are necessary to achieve high performance, with special emphasis on the scale-space: we introduce a twin pyramid for computing Difference-of-Gaussians, which enables a flexible center-surround ratio. The resulting system, called VOCUS2, is elegant and coherent in structure, fast, and computes saliency at the pixel level. It is not only suitable for images with few objects, but also for complex scenes as captured by mobile devices. Furthermore, we integrate the saliency system into an object proposal generation framework to obtain segment-based saliency maps and boost the results for salient object segmentation. We show that our system achieves state-of-the-art performance on a large collection of benchmark data. Simone Frintrop, Thomas Werner, Germán Martín García |
CVPR | 1 |
| 2015 | Saliency-based object discovery on RGB-D data with a late-fusion approachabstractWe present a novel method based on saliency and segmentation to generate generic object candidates from RGB-D data. Our method uses saliency as a cue to roughly estimate the location and extent of the objects present in the scene. Salient regions are used to glue together the segments obtained from over-segmenting the scene by either color or depth segmentation algorithms, or by a combination of both. We suggest a late-fusion approach that first extracts segments from color and depth independently before fusing them to exploit that the data is complementary. Furthermore, we investigate several mechanisms for ranking the object candidates. We evaluate on one publicly available dataset and on one challenging sequence with a high degree of clutter. The results show that we are able to retrieve most objects in real-world indoor scenes and clearly outperform other state-of-the art methods. Germán Martín García, Ekaterina Potapova, Thomas Werner, Michael Zillich, Markus Vincze, Simone Frintrop |
ICRA | 6 |
| 2015 | Sequence-level object candidates based on saliency for generic object recognition on mobile systemsabstractIn this paper, we propose a novel approach for generating generic object candidates for object discovery and recognition in continuous monocular video. Such candidates have recently become a popular alternative to exhaustive window-based search as basis for classification. Contrary to previous approaches, we address the candidate generation problem at the level of entire video sequences instead of at the single image level. We propose a processing pipeline that starts from individual region candidates and tracks them over time. This enables us to group candidates for similar objects and to automatically filter out inconsistent regions. For generating the per-frame candidates, we introduce a novel multi-scale saliency approach that achieves a higher per-frame recall with fewer candidates than current state-of-the-art methods. Taken together, those two components result in a significant reduction of the number of object candidates compared to frame level methods, while keeping a consistently high recall. Esther Horbert, Germán Martín García, Simone Frintrop, Bastian Leibe |
ICRA | 3 |
| 2015 | Saliency-Guided Object Candidates Based on Gestalt Principles
Thomas Werner, Germán Martín García, Simone Frintrop |
ICVS | 3 |
| 2014 | Workshop on attention models in robotics: visual systems for better HRIabstractAttention is a concept of human perception that enables human subjects to select the potentially relevant parts out of the huge amount of sensory data and that enables interactions with other human subjects by sharing attention with each other. These abilities are also of large interest for autonomous robots, therefore, interest in modeling concepts of human attention computationally has increased strongly in the robotics community during the last decade. Especially in human-robot interaction, the ability to detect what a human partner is attending to and to act in a similar way to enable intuitive communication, are important skills for a robotic system. Michael Zillich, Simone Frintrop, Fiora Pirri, Ekaterina Potapova, Markus Vincze |
HRI | 2 |
| 2014 | A Cognitive Approach for Object DiscoveryabstractObject discovery is the task of detecting unknown objects in images. The task is of large interest in many fields of machine vision, ranging from the automatic analysis of web images to interpreting data of a mobile robot or a driver assistant system. Here, we present a new approach for object discovery, based on findings of the human visual system. Proto-objects are detected with a segmentation module, generating perceptually coherent image regions. In parallel, a saliency system detects regions of interest in images and serves to select segments, depending on their saliency. We obtain very good results on a database of salient objects and on real-world office scenes. Simone Frintrop, Germán Martín García, Armin B. Cremers |
ICPR | 1 |
| 2014 | A Multisize Superpixel Approach for Salient Object Detection Based on Multivariate Normal Distribution EstimationabstractThis paper presents a new method for salient object detection based on a sophisticated appearance comparison of multisize superpixels. Those superpixels are modeled by multivariate normal distributions in CIE-Lab color space, which are estimated from the pixels they comprise. This fitting facilitates an efficient application of the Wasserstein distance on the Euclidean norm ( [Formula: see text]) to measure perceptual similarity between elements. Saliency is computed in two ways. On the one hand, we compute global saliency by probabilistically grouping visually similar superpixels into clusters and rate their compactness. On the other hand, we use the same distance measure to determine local center-surround contrasts between superpixels. Then, an innovative locally constrained random walk technique that considers local similarity between elements balances the saliency ratings inside probable objects and background. The results of our experiments show the robustness and efficiency of our approach against 11 recently published state-of-the-art saliency detection methods on five widely used benchmark data sets. Lei Zhu 0010, Dominik A. Klein, Simone Frintrop, Zhiguo Cao 0001, Armin B. Cremers |
IEEE Trans. Image Process. | 3 |
| 2013 | A Computational Framework for Attentional 3D Object Detection
Germán Martín García, Simone Frintrop |
CogSci | 2 |
| 2013 | Multi-scale region-based saliency detection using W2 distance on N-dimensional normal distributionsabstractWe present a new segment-based method for saliency detection based on multi-size superpixels that combines local and global saliency cues. We extract superpixels at several scales and represent each superpixel with a normal distribution in CIE-Lab space estimated from its associated pixels. Global saliency is computed by grouping similar superpixels to estimate the spatial distribution of colors, while local saliency detection is achieved by determining the center-surround contrast of neighboring superpixels. Both methods rely on the Wasserstein distance on L2norm (W2) to measure perceptual (dis-)similarity between superpixels. Additionally, we propose a Saliency Flow technique to refine the local saliency map. Our approach uses very few empirical parameters and outperforms 6 recent state-of-the-art saliency detection methods in terms of several evaluations on a widely used benchmark. Lei Zhu 0010, Dominik A. Klein, Simone Frintrop, Zhiguo Cao 0001, Armin B. Cremers |
ICIP | 3 |
| 2011 | Center-surround divergence of feature statistics for salient object detectionabstractIn this paper, we introduce a new method to detect salient objects in images. The approach is based on the standard structure of cognitive visual attention models, but realizes the computation of saliency in each feature dimension in an information-theoretic way. The method allows a consistent computation of all feature channels and a well-founded fusion of these channels to a saliency map. Our framework enables the computation of arbitrarily scaled features and local center-surround pairs in an efficient manner. We show that our approach outperforms eight state-of-the-art saliency detectors in terms of precision and recall. Dominik A. Klein, Simone Frintrop |
ICCV | 2 |
| 2010 | Learning context-based feature descriptors for object trackingabstractA major problem with previous object tracking approaches is adapting object representations depending on scene context to account for changes in illumination, viewpoint changes, etc. To adapt our previous approach to deal with background changes, here we first derive some clusters from a training sequence and the corresponding object representations for those clusters. Next, for each frame of a separate test sequence, its nearest background cluster is determined and then the corresponding descriptor of that cluster is used for object representation in this frame. Experiments show that the proposed approach tracks objects and persons in natural scenes more effectively. Ali Borji, Simone Frintrop |
HRI | 2 |
| 2010 | General object tracking with a component-based target descriptorabstractIn this paper, we present a component-based visual object tracker for mobile platforms. The core of the technique is a component-based descriptor that captures the structure and appearance of a target in a flexible way. This descriptor can be learned quickly from a single training image and is easily adaptable to different objects. The descriptor is integrated into the observation model of a visual tracker based on the well known Condensation algorithm. We show that the approach is applicable to a large variety of objects and in different environments with cluttered backgrounds and a moving camera. The method is robust to illumination and viewpoint changes and applicable to indoor as well as outdoor scenes. Simone Frintrop |
ICRA | 1 |
| 2010 | Visual landmark generation and redetection with a single feature per frameabstractIn this paper we show that visual landmark generation and redetection is possible with a single feature per frame. The approach is based on the assumption that highly discriminative regions are easily redetectable in subsequent frames as well as in frames visited from different viewpoints. We investigate which feature detectors fit for this purpose and under which conditions the discriminability applies. The approach is tested in a topological localization scenario in which the best feature is tracked over several frames to build landmarks. We show that we can represent a large environment with a few salient landmarks and that a large percentage of these landmarks is robustly redetectable from different viewpoints. Simone Frintrop, Armin B. Cremers |
ICRA | 1 |
| 2010 | Adaptive real-time video-tracking for arbitrary objectsabstractIn this paper, we present a visual object tracker for mobile systems that is able to specialize to individual objects during tracking. The core of our method is a novel observation model and the way it is automatically adapted to a changing object and background appearance over time. The model is integrated into the well known Condensation algorithm (SIR filter) for statistical inference, and it consists of a boosted ensemble of simple threshold classifiers built upon center-surround Haar-like features, which the filter continuously updates based on the images perceived. We present optimizations and reasonable approximations to limit the computational costs. Thus, the final algorithms are capable of processing video input at real-time. To experimentally investigate the gain of adapting the observation model we compare two different approaches with a non-adapting version of our observation model: maintaining a single observation model for all particles, and maintaining individual observation models for each particle. In addition, experiments were conducted to compare system performances between the proposed algorithms and two other state of the art Condensation based tracking approaches. Dominik A. Klein, Dirk Schulz 0001, Simone Frintrop, Armin B. Cremers |
IROS | 3 |
| 2010 | Computational visual attention systems and their cognitive foundations: A surveyabstractBased on concepts of the human visual system, computational visual attention systems aim to detect regions of interest in images. Psychologists, neurobiologists, and computer scientists have investigated visual attention thoroughly during the last decades and profited considerably from each other. However, the interdisciplinarity of the topic holds not only benefits but also difficulties: Concepts of other fields are usually hard to access due to differences in vocabulary and lack of knowledge of the relevant literature. This article aims to bridge this gap and bring together concepts and ideas from the different research areas. It provides an extensive survey of the grounding psychological and biological research on visual attention as well as the current state of the art of computational systems. Furthermore, it presents a broad range of applications of computational attention systems in fields like computer vision, cognitive systems, and mobile robotics. We conclude with a discussion on the limitations and open questions in the field. Simone Frintrop, Erich Rome, Henrik I. Christensen |
ACM Trans. Appl. Percept. | 1 |
| 2009 | Most salient region trackingabstractIn this paper, we introduce a cognitive approach for object tracking from a mobile platform. The approach is based on a biologically motivated attention system which is able to detect regions of interest in images based on concepts of the human visual system. A top-down guided visual search module of the system enables to especially favor features which fit to a previously learned target object. Here, the appearance of an object is learned online within the first image in which it is detected. In subsequent images, the attention system searches for the target features and builds a top-down, target-related saliency map. This enables to focus on the most relevant features of especially this object in especially this scene without knowing anything about a particular object model or scene in advance. The system is able to operate in real-time and to cope with the requirements of real-world tasks such as illumination variations and other moving objects. Simone Frintrop, Markus Kessel |
ICRA | 1 |
| 2009 | Boosting with a Joint Feature Pool from Different Sensors
Dominik A. Klein, Dirk Schulz 0001, Simone Frintrop |
ICVS | 3 |
| 2008 | Active gaze control for attentional visual SLAMabstractIn this paper, we introduce an approach to active camera control for visual SLAM. Features, detected by a biologically motivated attention system, are tracked over several frames to determine stable landmarks. Matching of features to database entries enables global loop closing. The focus of this paper is the active camera control module, which supports the system with three behaviours: (i) A tracking behaviour tracks promising landmarks and prevents them from leaving the field of view, (ii) A redetection behaviour directs the camera actively to regions where landmarks are expected and thus supports loop closing, (iii) Finally, an exploration behaviour investigates regions without landmarks and enables a more uniform distribution of landmarks. Several real-world experiments show that the active camera control outperforms the passive system considerably. Simone Frintrop, Patric Jensfelt |
ICRA | 1 |
| 2008 | Attentional Landmarks and Active Gaze Control for Visual SLAMabstractThis paper is centered around landmark detection, tracking, and matching for visual simultaneous localization and mapping using a monocular vision system with active gaze control. We present a system that specializes in creating and maintaining a sparse set of landmarks based on a biologically motivated feature-selection strategy. A visual attention system detects salient features that are highly discriminative and ideal candidates for visual landmarks that are easy to redetect. Features are tracked over several frames to determine stable landmarks and to estimate their 3-D position in the environment. Matching of current landmarks to database entries enables loop closing. Active gaze control allows us to overcome some of the limitations of using a monocular vision system with a relatively small field of view. It supports 1) the tracking of landmarks that enable a better pose estimation, 2) the exploration of regions without landmarks to obtain a better distribution of landmarks in the environment, and 3) the active redetection of landmarks to enable loop closing in situations in which a fixed camera fails to close the loop. Several real-world experiments show that accurate pose estimation is obtained with the presented system and that active camera control outperforms the passive approach. Simone Frintrop, Patric Jensfelt |
IEEE Trans. Robotics | 1 |
| 2006 | Attentional Landmark Selection for Visual SLAMabstractIn this paper, we introduce a new method to automatically detect useful landmarks for visual SLAM. A biologically motivated attention system detects regions of interest which "pop-out" automatically due to strong contrasts and the uniqueness of features. This property makes the regions easily redetectable and thus they are useful candidates for visual landmarks. Matching based on scene prediction and feature similarity allows not only short-term tracking of the regions, but also redetection in loop closing situations. The paper demonstrates how regions are determined and how they are matched reliably. Various experimental results on real-world data show that the landmarks are useful with respect to be tracked in consecutive frames and to enable closing loops Simone Frintrop, Patric Jensfelt, Henrik I. Christensen |
IROS | 1 |
| 2005 | Robust Object Detection at Regions of Interest with an Application in Ball RecognitionabstractIn this paper, we present a new combination of a biologically inspired attention system (VOCUS – Visual Object detection with a CompUtational attention System) with a robust object detection method. As an application, we built a reliable system for ball recognition in the RoboCup context. Firstly, VOCUS finds regions of interest generating a hypothesis for possible locations of the ball. Secondly, a fast classifier verifies the hypothesis by detecting balls at regions of interest. The combination of both approaches makes the system highly robust and eliminates false detections. Furthermore, the system is quickly adaptable to balls in different scenarios: The complex classifier is universally applicable to balls in every context and the attention system improves the performance by learning scenario-specific features quickly from only a few training examples. Sara Mitri, Simone Frintrop, Kai Pervölz, Hartmut Surmann, Andreas Nüchter |
ICRA | 2 |
| 2005 | A Bimodal Laser-Based Attention System
Simone Frintrop, Erich Rome, Andreas Nüchter, Hartmut Surmann |
Comput. Vis. Image Underst. | 1 |
| 2004 | Saliency-based object recognition in 3D dataabstractThis paper presents a robust and real-time capable recognition system for the fast detection and classification of objects in spatial 3D data. Depth and reflection data from a 3D laser scanner are rendered into images and fed into a saliency-based visual attention system that detects regions of potential interest. Only these regions are examined by a fast classifier. The time saving of classifying objects in salient regions rather than in complete images is linear with the number of trained object classes. Robustness is achieved by the fusion of the bi-modal scanner data; in contrast to camera images, this data is completely illumination independent. The recognition system is trained for two different object classes and evaluated on real indoor data. Simone Frintrop, Andreas Nüchter, Hartmut Surmann, Joachim Hertzberg |
IROS | 1 |
| 2003 | An Attentive, Multi-modal Laser "Eye"
Simone Frintrop, Erich Rome, Andreas Nüchter, Hartmut Surmann |
ICVS | 1 |
| 2001 | Robust Localization Using Context in Omnidirectional ImagingabstractThis work presents the concept to recover and utilize the visual context in panoramic images. Omnidirectional imaging has become recently an efficient basis for robot navigation. The proposed Bayesian reasoning over local image appearances enables to reject false hypotheses which do not fit the structural constraints in corresponding feature trajectories. The methodology is proved with real image data from an office robot to dramatically increase the localization performance in the presence of severe occlusion effects, particularly in noisy environments, and to recover rotational information on the fly. Lucas Paletta, Simone Frintrop, Joachim Hertzberg |
ICRA | 2 |