Gunther Heidemann

dblp:14/1385 · DBLP profile ↗
← Back
59ranked-venue papers
18as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 44 · 15 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-authorSystems, architecture and hardware · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Information extraction and text analysis · 36% Trustworthy machine learning · 36% Robot manipulation · 11%
Computer graphics and multimedia
3 papers
Visualization and visual analytics · 46% Multimedia analysis and retrieval · 30% Image and video processing · 17%

Topics — the 14 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
0.812024
FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense Disambiguation · EMNLP 2024
Natural language and speech › Information extraction and text analysis
word sense disambiguation
0.812024
FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense Disambiguation · EMNLP 2024
Visualization and visual analytics › movement data analysis
trajectory clustering
0.212013
Interactive Schematic Summaries for Faceted Exploration of Surveillance Video · IEEE Trans. Multim. 2013
Multimedia analysis and retrieval › multimedia browsing
video exploration
0.212013
Interactive Schematic Summaries for Faceted Exploration of Surveillance Video · IEEE Trans. Multim. 2013
Computer vision › Video understanding and tracking
background subtraction
0.112011
Evaluation of background subtraction techniques for video surveillance · CVPR 2011
Robotics › Robot manipulation
tactile sensing
0.122007
Acquisition and Application of a Tactile Database · ICRA 2007
Dynamic Tactile Sensing for Object Identification · ICRA 2004
Robotics › Robot manipulation › tactile sensing
tactile object recognition
0.112007
Acquisition and Application of a Tactile Database · ICRA 2007
Image and video processing › image statistics › statistical image modeling
natural image statistics
0.112006
The Principal Components of Natural Images Revisited · IEEE Trans. Pattern Anal. Mach. Intell. 2006
Computer vision › Image recognition and object detection
object recognition
0.122004
Dynamic Tactile Sensing for Object Identification · ICRA 2004
Focus-of-Attention from Local Color Symmetries · IEEE Trans. Pattern Anal. Mach. Intell. 2004
Computer vision › 3D vision › low-level vision
feature detection
0.112005
The long-range saliency of edge- and corner-based salient points · IEEE Trans. Image Process. 2005
Multimedia analysis and retrieval
video surveillance
0.012013
Interactive Schematic Summaries for Faceted Exploration of Surveillance Video · IEEE Trans. Multim. 2013
Robotics › Robot manipulation › object perception
object identification
0.012004
Dynamic Tactile Sensing for Object Identification · ICRA 2004
Computer vision › Video understanding and tracking
video surveillance
0.012011
Evaluation of background subtraction techniques for video surveillance · CVPR 2011
Haptics and multimodal interaction
tactile perception
0.012004
Dynamic Tactile Sensing for Object Identification · ICRA 2004

Methods — techniques the papers use, named apart from their topics

adversarial dataset · 0.8trajectory bundling · 0.2interactive clustering · 0.2evaluation · 0.1benchmarking · 0.1principal component analysis · 0.1time series classification · 0.1neural architecture · 0.1local color symmetry measure · 0.1trajectory calculation · 0.1neural network classification · 0.1edge-based salient points · 0.1corner-based salient points · 0.1
YearPublicationVenuePosition
2024 FOOL ME IF YOU CAN! An Adversarial Dataset to Investigate the Robustness of LMs in Word Sense Disambiguation
abstract
Mohamad Ballout, Anne Dedert, Nohayr Muhammad Abdelmoneim, Ulf Krumnack, Gunther Heidemann, Kai-Uwe Kühnberger. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Mohamad Ballout, Anne Dedert, Nohayr Abdelmoneim, Ulf Krumnack, Gunther Heidemann, Kai-Uwe Kühnberger
EMNLP5
2024 Efficient Knowledge Distillation: Empowering Small Language Models with Teacher Model Insights
Mohamad Ballout, Ulf Krumnack, Gunther Heidemann, Kai-Uwe Kühnberger
NLDB (1)3
2024 Enhancing Small Language Models via ChatGPT and Dataset Augmentation
Tom Pieper, Mohamad Ballout, Ulf Krumnack, Gunther Heidemann, Kai-Uwe Kühnberger
NLDB (2)4
2023 Show Me How It's Done: The Role of Explanations in Fine-Tuning Language Models
Mohamad Ballout, Ulf Krumnack, Gunther Heidemann, Kai-Uwe Kühnberger
ACML3
2020 Hierarchical Modeling with Neurodynamical Agglomerative Analysis
Michael Marino, Georg Schröter, Gunther Heidemann, Joachim Hertzberg
ICANN (1)3
2018 Object of Interest Segmentation in Video Sequences with Gaze Data
abstract
Object of Interest (OoI) segmentation in video sequences is, due to its temporal and spatial complexity, still a difficult task. Thus, it cannot be done automatically, and interactive segmentation methods are inconvenient and time-consuming. Overcoming these drawbacks and performing OoI segmentation in real-time, this paper introduces a new approach for gaze-based OoI segmentation. The user fixates on the OoI throughout watching a video sequence. The users' gaze provides relevant information on the location of the OoI, which is then used for segmenting the OoI. In comparison to present segmentation methods, our approach notably reduced user-interactions required to attain pixel-accurate segmentations. The first evaluation of our gaze-based OoI segmentation method shows promising results. These are close to, and sometimes even exceeding the quality of the polygon-based ground truth segmentation. In conclusion, these findings encourage further exploration of gaze-based OoI segmentation.
Laura Krieger, Gunther Heidemann, Julius Schöning
IPAS2
2018 Image-based 3D Reconstruction: Neural Networks vs. Multiview Geometry
abstract
Methods using multiple view geometry (MVG), like Structure from Motion (SfM), are still the dominant approaches for image-based 3D reconstruction. These reconstruction methods have become quite robust and accurate. However, how robust and accurate can artificial neural networks (ANNs) reconstruct a priori unknown 3D objects? Exceed the restriction of object categories this paper evaluates ANNs for reconstructing arbitrary 3D objects. With the use of a synthetic scalable cube dataset for training, testing and validating ANNs, it is shown that ANNs are capable of learning mathematical principles of 3D reconstruction. As inspiration for the design of the different ANNs architectures, the global, hierarchical, and incremental key-point matching strategies of SfM approaches were taken into account. Based on these benchmarks and a review of the used dataset, it is shown that voxel-based 3D reconstruction cannot be scaled. Thus, voxel-based reconstruction might be misleading for capturing the complexity of real-world images. Also, the benchmark results show that ANNs have the same limitation by reconstructing unknown object categories as current MVG approach.
Julius Schöning, Gunther Heidemann
IPAS2
2018 An Empirical Study on Bidirectional Recurrent Neural Networks for Human Motion Recognition
abstract
The deep recurrent neural networks (RNNs) and their associated gated neurons, such as Long Short-Term Memory (LSTM) have demonstrated a continued and growing success rates with researches in various sequential data processing applications, especially when applied to speech recognition and language modeling. Despite this, amongst current researches, there are limited studies on the deep RNNs architectures and their effects being applied to other application domains. In this paper, we evaluated the different strategies available to construct bidirectional recurrent neural networks (BRNNs) applying Gated Recurrent Units (GRUs), as well as investigating a reservoir computing RNNs, i.e., Echo state networks (ESN) and a few other conventional machine learning techniques for skeleton-based human motion recognition. The evaluation of tasks focuses on the generalization of different approaches by employing arbitrary untrained viewpoints, combined together with previously unseen subjects. Moreover, we extended the test by lowering the subsampling frame rates to examine the robustness of the algorithms being employed against the varying of movement speed.
Pattreeya Tanisaro, Gunther Heidemann
TIME2
2017 Classifying Bio-Inspired Model of Point-Light Human Motion Using Echo State Networks
Pattreeya Tanisaro, Constantin Lehman, Leon René Sütfeld, Gordon Pipa, Gunther Heidemann
ICANN (1)5
2017 Providing Video Annotations in Multimedia Containers for Visualization and Research
abstract
There is an ever increasing amount of video data sets which comprise additional metadata, such as object labels, tagged events, or gaze data. Unfortunately, metadata are usually stored in separate files in custom-made data formats, which reduces accessibility even for experts and makes the data inaccessible for non-experts. Consequently, we still lack interfaces for many common use cases, such as visualization, streaming, data analysis, machine learning, high-level understanding and semantic web integration. To bridge this gap, we want to promote the use of existing multimedia container formats to establish a standardized method of incorporating content and metadata. This will facilitate visualization in standard multimedia players, streaming via the Internet, and easy use without conversion, as shown in the attached demonstration video and files. In two prototype implementations, we embed object labels, gaze data from eye-tracking and the corresponding video into a single multimedia container and visualize this data using a media player. Based on this prototype, we discuss the benefit of our approach as a possible standard. Finally, we argue for the inclusion of MPEG-7 in multimedia containers as a further improvement.
Julius Schöning, Patrick Faion, Gunther Heidemann, Ulf Krumnack
WACV3
2016 Time Series Classification Using Time Warping Invariant Echo State Networks
abstract
For many years, neural networks have gained gigantic interest and their popularity is likely to continue because of the success stories of deep learning. Nonetheless, their applications are mostly limited to static and not temporal patterns. In this paper, we apply time warping invariant Echo State Networks (ESNs) to time-series classification tasks using datasets from various studies in the UCR archive. We also investigate the influence of ESN architecture and spectral radius of the network in view of general characteristics of data, such as dataset type, number of classes, and amount of training data. We evaluate our results comparing it to other state-of-the-art methods, using One Nearest Neighbor (1-NN) with Euclidean Distance (ED), Dynamic Time Warping (DTW) and best warping window DTW.
Pattreeya Tanisaro, Gunther Heidemann
ICMLA2
2016 Pixel-wise Ground Truth Annotation in Videos - An Semi-automatic Approach for Pixel-wise and Semantic Object Annotation
abstract
In the last decades, a large diversity of automatic, semi-automatic and manual approaches for video segmentation and knowledge extraction from video-data has been proposed. Due to the high complexity in both the spatial and temporal domain, it continues to be a challenging research area. In order to develop, train, and evaluate new algorithms, ground truth of video-data is crucial. Pixel-wise annotation of ground truth is usually time-consuming, does not contain semantic relations between objects and uses only simple geometric primitives. We provide a brief review of related tools for video annotation, and introduce our novel interactive and semi-automatic segmentation tool iSeg. Extending an earlier implementation, we improved iSeg with a semantic time line, multithreading and the use of ORB features. A performance evaluation of iSeg on four data sets is presented. Finally, we discuss possible opportunities and applications of semantic polygon-shaped video annotation, such as 3D reconstruction and video inpainting.
Julius Schöning, Patrick Faion, Gunther Heidemann
ICPRAM3
2015 Evaluation of Multi-view 3D Reconstruction Software
Julius Schöning, Gunther Heidemann
CAIP (2)2
2015 FOREST - A Flexible Object Recognition System
Julia Möhrmann, Gunther Heidemann
ICPRAM (2)2
2015 Interactive 3D Modeling - A Survey-based Perspective on Interactive 3D Reconstruction
Julius Schöning, Gunther Heidemann
ICPRAM (2)2
2015 Semi-automatic ground truth annotation in videos: An interactive tool for polygon-based object annotation and segmentation
abstract
Knowledge extraction from video data is challenging due to its high complexity in both the spatial and temporal domain. Ground truth is crucial for the evaluation and the adaptation of algorithms to new domains. Unfortunately, ground truth annotation is inconvenient and time consuming. Common annotation tools mostly rely on simple geometric primitives such as rectangles or ellipses. Here we propose a novel, interactive and semi-automatic process, which actively asks for user input if the result of the automatic annotation appears to be incorrect. After a brief review of related tools for video annotation, we explain our proposed semi-automatic method iSeg using a prototype implementation. iSeg has been tested on two visual stimulus datasets for eye tracking experiments and on two surveillance datasets. The experimental results and the usability are compared to existing annotation tools. Finally, we discuss the properties and opportunities of polygon-based video annotation.
Julius Schöning, Patrick Faion, Gunther Heidemann
K-CAP3
2013 Semi-automatic Image Annotation
Julia Möhrmann, Gunther Heidemann
CAIP (2)2
2013 Interactive Schematic Summaries for Faceted Exploration of Surveillance Video
abstract
We present a scalable technique to explore surveillance videos by scatter/gather browsing of trajectories of moving objects. Trajectories are clustered according to a variety of properties, such as location, orientation, and velocity that can be selected by the users. These properties allow for faceted video exploration and refinement of previous browsing steps. The proposed approach facilitates interactive clustering of trajectories by an effective way of cluster visualization that we term schematic summaries. This novel visualization illustrates cluster summaries in a schematic, nonphotorealistic style. To reduce visual clutter, we introduce the trajectory bundling technique. Further, schematic summaries include a timeline view and a showcase view to represent the facets present in a cluster. The fusion of schematic summaries, a variety of facets, and user interaction lead to efficient hierarchical exploration of video data. Examples of different browsing scenarios and initial user feedback demonstrate the potentials of our method.
Markus Höferlin, Benjamin Höferlin, Gunther Heidemann, Daniel Weiskopf
IEEE Trans. Multim.3
2012 Learning a Visual Attention Model for Adaptive Fast-forward in Video Surveillance
Benjamin Höferlin, Hermann Pflüger, Markus Höferlin, Gunther Heidemann, Daniel Weiskopf
ICPRAM (2)4
2012 State of the Art Report on Video-Based Graphics and Video Visualization
abstract
Abstract In recent years, a collection of new techniques which deal with video as input data, emerged in computer graphics and visualization. In this survey, we report the state of the art in video‐based graphics and video visualization. We provide a review of techniques for making photo‐realistic or artistic computer‐generated imagery from videos, as well as methods for creating summary and/or abstract visual representations to reveal important features and events in videos. We provide a new taxonomy to categorize the concepts and techniques in this newly emerged body of knowledge. To support this review, we also give a concise overview of the major advances in automated video analysis, as some techniques in this field (e.g. feature extraction, detection, tracking and so on) have been featured in video‐based modelling and rendering pipelines for graphics and visualization.
Rita Borgo, Min Chen 0001, Ben Daubney, Edward Grundy, Gunther Heidemann, Benjamin Höferlin, Markus Höferlin, Heike Leitte, Daniel Weiskopf, Xianghua Xie
Comput. Graph. Forum5
2012 Evaluation of Fast-Forward Video Visualization
abstract
We evaluate and compare video visualization techniques based on fast-forward. A controlled laboratory user study (n = 24) was conducted to determine the trade-off between support of object identification and motion perception, two properties that have to be considered when choosing a particular fast-forward visualization. We compare four different visualizations: two representing the state-of-the-art and two new variants of visualization introduced in this paper. The two state-of-the-art methods we consider are frame-skipping and temporal blending of successive frames. Our object trail visualization leverages a combination of frame-skipping and temporal blending, whereas predictive trajectory visualization supports motion perception by augmenting the video frames with an arrow that indicates the future object trajectory. Our hypothesis was that each of the state-of-the-art methods satisfies just one of the goals: support of object identification or motion perception. Thus, they represent both ends of the visualization design. The key findings of the evaluation are that object trail visualization supports object identification, whereas predictive trajectory visualization is most useful for motion perception. However, frame-skipping surprisingly exhibits reasonable performance for both tasks. Furthermore, we evaluate the subjective performance of three different playback speed visualizations for adaptive fast-forward, a subdomain of video fast-forward.
Markus Höferlin, Kuno Kurzhals, Benjamin Höferlin, Gunther Heidemann, Daniel Weiskopf
IEEE Trans. Vis. Comput. Graph.4
2011 Evaluation of background subtraction techniques for video surveillance
abstract
Background subtraction is one of the key techniques for automatic video analysis, especially in the domain of video surveillance. Although its importance, evaluations of recent background subtraction methods with respect to the challenges of video surveillance suffer from various shortcomings. To address this issue, we first identify the main challenges of background subtraction in the field of video surveillance. We then compare the performance of nine background subtraction methods with post-processing according to their ability to meet those challenges. Therefore, we introduce a new evaluation data set with accurate ground truth annotations and shadow masks. This enables us to provide precise in-depth evaluation of the strengths and drawbacks of background subtraction methods.
Sebastian Brutzer, Benjamin Höferlin, Gunther Heidemann
CVPR3
2011 Similarity Calculation with Length Delimiting Dictionary Distance
abstract
The Normalized Compression Distance (NCD) has gained considerable interest in pattern recognition as a similarity measure applicable to unstructured data of very different domains, such as text, DNA sequences, or images. NCD uses existing compression programs such as gzip to compute similarity between objects. NCD has unique features: It does not require any prior knowledge, data preprocessing, feature extraction, domain adaptation or any parameter settings. Further, the NCD can be applied to symbolic data and raw signals alike. In this paper we decompose the NCD and introduce a method to measure compression-based similarity without the need to use compression. The Length Delimiting Dictionary Distance (LD3) takes the one component essential in compression methods, the dictionary generation, and strips the NCD of all dispensable components. The LD3performs "compression based pattern recognition without compression", keeping all of the above benefits of the NCD while achieving better speed and recognition rates. We first review the NCD, introduce LD3as the "essence" of NCD, and evaluate the LD3based on language tree experiments, authorship recognition, and genome phylogeny data.
Andre Burkovski, Sebastian Klenk, Gunther Heidemann
ICTAI3
2011 Efficient monocular vehicle orientation estimation using a tree-based classifier
abstract
For automotive assistance systems, on-road vehicle detection is a key challenge to forward collision warning. Along with detecting existence, determining a vehicle's orientation plays an important role in correctly predicting maneuvers. In this paper, an approach to remotely estimate vehicle orientations from monocular images is presented. The proposed system operates on a per frame basis and does not require any depth cues. Orientation estimation is performed by analyzing the position of the vehicle's rear section relative to the overall vehicle outline. Both position types are determined using a newly devised tree-structured classifier. Based on the cascaded structure by Viola and Jones, the pro posed classifier adapts itself to the problem's structure, dividing the overall problem into parts that require fewer weak learners to solve. To find partitions that simplify the classification task, a quality criterion measuring class separability is optimized using the Simulated Annealing algorithm. To further increase processing speed, the number of tree nodes to be traversed is drastically reduced by a two-staged boosting procedure, training a classifier that decides which branch to take. Experiments show the relevance and effectiveness of the proposed concepts.
Michael Gabb, Otto Löhlein, Matthias F. Oberländer, Gunther Heidemann
Intelligent Vehicles Symposium4
2011 Interactive schematic summaries for exploration of surveillance video
abstract
We present a new and scalable technique to explore surveillance videos by scatter/gather browsing of trajectories of moving objects. The proposed approach facilitates interactive clustering of trajectories by an effective way of cluster visualization that we term schematic summaries. This novel visualization illustrates cluster summaries in a schematic, non-photorealistic style. To reduce visual clutter, we introduce the trajectory bundling technique. The fusion of schematic summaries and user interaction leads to efficient hierarchical exploration of video data. Examples of different browsing scenarios demonstrate the effectiveness of the proposed method.
Markus Höferlin, Benjamin Höferlin, Daniel Weiskopf, Gunther Heidemann
ICMR4
2011 Information-based adaptive fast-forward for visual surveillance
Benjamin Höferlin, Markus Höferlin, Daniel Weiskopf, Gunther Heidemann
Multim. Tools Appl.4
2010 Relational feature engineering of natural language processing
abstract
We present a new framework for feature engineering of natural language processing that is based on a relational data model of text. It includes fast and flexible methods for implementing and extracting new features and thereby reduces the effort of creating an NLP system for a particular task.
Hamidreza Kobdani, Hinrich Schütze, Andre Burkovski, Wiltrud Kessler, Gunther Heidemann
CIKM5
2010 Robust 1D Barcode Recognition on Mobile Devices
abstract
In the following we will describe a novel method for decoding linear barcodes from blurry camera images. Our goal was to develop a algorithm that can be used on mobile devices to recognize product numbers from EAN or UPC barcodes.
Johann C. Rocholl, Sebastian Klenk, Gunther Heidemann
ICPR3
2008 A Spatio-temporal Extension of the SUSAN-Filter
Benedikt Kaiser, Gunther Heidemann
ICANN (1)2
2008 Interactive feature visualization for image retrieval
abstract
Most systems for content based image retrieval (CBIR) employ low level image features as a similarity measure. The problem of CBIR systems is that they are a ldquoblack boxrdquo to the user: Queries are specified by sample images, but the features which the CBIR system actually uses are unknown to the user. Hence, unexpected results are difficult to interpret. The problem becomes worse for inexperienced users, who expect the system to understand their query on a symbolic level, while in reality the CBIR system just extracts close-to-signal features. Here we propose to make CBIR systems more ldquotransparentrdquo by visualization of the employed features. Since non-experts should be able to operate the CBIR system, we argue that features should be visualized as prototypical, artificial images, rather than feature-specific visualizations (such as bar-diagrams for a histogram). We present the visualization of two widely used feature classes, color histograms and texture features, and evaluate in a user study how well the visualizations can be interpreted.
Johannes Imo, Sebastian Klenk, Gunther Heidemann
ICPR3
2008 Qualitative analysis of spatio-temporal event detectors
abstract
Interest point detection is an established method to select relevent image regions. Such techniques use features like corners or edges, which are known to indicate regions likely to hold patterns of interest. Selection of such regions increases processing efficiency. For the recognition of motion, however, such context-free methods are still very rare. Though there are numerous methods to find space-time volumes of motion in image sequences, most aim at finding just motion as a such, not volumes which are more promising for analysis than others. Therefore Laptev and Lindeberg (2005) generalized the Harris detector to the spatio-temporal domain. But the problem remains to evaluate what kind of motion is captured by a detector. For example, the detector of Laptev and Lindeberg should capture ¿corners¿ - like the original 2D-version of Harris and Stephens (1988) - but what does that mean for motion? Therefore we present an approach to visualize events which were selected by a spatio-temporal interest point detector. Since the analysis of single examples is not fruitful, we use clustering to analyze large quantities of space-time volumes selected by a detector. The resulting cluster centers are prototypical events, representing the types of events the detector responds to. Thus a qualitative yet statistically exhaustive analysis of detector properties is possible.
Benedikt Kaiser, Gunther Heidemann
ICPR2
2008 The visual active memory perspective on integrated recognition systems
Christian Bauckhage, Sven Wachsmuth, Marc Hanheide, Sebastian Wrede 0001, Gerhard Sagerer, Gunther Heidemann, Helge J. Ritter
Image Vis. Comput.6
2008 Region saliency as a measure for colour segmentation stability
Gunther Heidemann
Image Vis. Comput.1
2007 Acquisition and Application of a Tactile Database
abstract
We present a database of 2D pressure profile time series as a testbed for tactile object and surface recognition. The tactile database captures the surfaces of household and toy objects by moving a 2D pressure sensor mounted to an industrial robot arm around the objects using real-time trajectory calculation. Thus, it represents different "views" of the objects in a similar way as the well known Columbia Object Image Library (COIL) captures different views of an object by a camera. As a first application, objects in the database are classified using a neural network architecture.
Matthias Schöpfer, Helge J. Ritter, Gunther Heidemann
ICRA3
2006 The Principal Components of Natural Images Revisited
abstract
This paper investigates the principal components (PCs) of natural gray and color images. A horizontal and vertical typology of PCs is found which leads to the identification of groups of basis functions for steerable bandpass filters. Using this system, the contribution of spatio-chromatic structure to the total variance can be quantified for selected spatial frequencies.
Gunther Heidemann
IEEE Trans. Pattern Anal. Mach. Intell.1
2005 Face detection and identification using a hierarchical feed-forward recognition architecture
abstract
We apply a hierarchical feed-forward neural architecture to the problem of face recognition. The network is similar to the neocognitron-approach and a two-layer variation of this architecture, which has previously been successfully applied to patch classification tasks. We extend this architecture to a three-layer one, which allows not only identification of image patches, but also detection in larger images. In the research area of face recognition, a lot of expertise has been developed for the problem of either identification or detection, but approaches which deal with both problems simultaneously are rarely to be found. In this work, we apply the hierarchical approach to this problem and evaluate the performance on artificial datasets.
Ingo Bax, Gunther Heidemann, Helge J. Ritter
IJCNN2
2005 SOM based image data structuring in an augmented reality scenario
abstract
Our research focuses on the development of a mobile augmented reality system which is capable of acquiring image data in an unrestricted environment and which provides a comfortable facility to label this data. To structure the image data modified MPEG-7 features are computed and by means of self organizing maps (SOM), the imagery can be labeled stepwise. First, the complete data set is projected onto the SOM using a combination of color and edge features. In a second step, selected parts of the imagery are retrained weighting the feature blocks depending on characteristics of the acquired image data. Within few steps, the partitioning leads to SOM nodes on which the projected imagery can be labeled as objects or rejected.
Holger Bekel, Gunther Heidemann, Helge J. Ritter
IJCNN2
2005 Unsupervised image categorization
Gunther Heidemann
Image Vis. Comput.1
2005 Interactive image data labeling using self-organizing maps in an augmented reality scenario
Holger Bekel, Gunther Heidemann, Helge J. Ritter
Neural Networks2
2005 The long-range saliency of edge- and corner-based salient points
abstract
A major goal of salient-point (SP) detection is increasing computational efficiency. Therefore, methods which can detect saliency of a large region by evaluation of only a small local patch are of particular interest. This paper checks for well-known detectors whether saliency outreaches the actual SPs.
Gunther Heidemann
IEEE Trans. Image Process.1
2004 Multimodal interaction in an augmented reality scenario
abstract
We describe an augmented reality system designed for online acquisition of visual knowledge and retrieval of memorized objects. The system relies on a head mounted camera and display, which allow the user to view the environment together with overlaid augmentations by the system. In this setup, communication by hand gestures and speech is mandatory as common input devices like mouse and keyboard are not available. Using gesture and speech, basically three types of tasks must be handled: (i) Communication with the system about the environment, in particular, directing attention towards objects and commanding the memorization of sample views; (ii) control of system operation, e.g. switching between display modes; and (iii) re-adaptation of the interface itself in case communication becomes unreliable due to changes in external factors, such as illumination conditions. We present an architecture to manage these tasks and describe and evaluate several of its key elements, including modules for pointing gesture recognition, menu control based on gesture and speech, and control strategies to cope with situations when vision becomes unreliable and has to be re-adapted by speech.
Gunther Heidemann, Ingo Bax, Holger Bekel
ICMI1
2004 Dynamic Tactile Sensing for Object Identification
abstract
We propose a neural architecture for the recognition of objects by haptics. We demonstrate its performance for a set of household objects and toys using a low cost 2D pressure sensor of coarse resolution, which is moved by a robot arm guided by contact points. The approach transfers the well known view-based method from computer vision to the domain of tactile sensing. However, in contrast to computer vision, not static frames but entire time series of 2D pressure profiles are evaluated.
Gunther Heidemann, Matthias Schöpfer
ICRA1
2004 Combining spatial and colour information for content based image retrieval
Gunther Heidemann
Comput. Vis. Image Underst.1
2004 Integrating context-free and context-dependent attentional mechanisms for gestural object reference
Gunther Heidemann, Robert Rae, Holger Bekel, Ingo Bax, Helge J. Ritter
Mach. Vis. Appl.1
2004 Focus-of-Attention from Local Color Symmetries
abstract
In this paper, a continuous valued measure for local color symmetry is introduced. The new algorithm is an extension of the successful gray value-based symmetry map proposed by Reisfeld et al. The use of color facilitates the detection of focus points (FPs) on objects that are difficult to detect using gray-value contrast only. The detection of FPs is aimed at guiding the attention of an object recognition system; therefore, FPs have to fulfill three major requirements: stability, distinctiveness, and usability. The proposed algorithm is evaluated for these criteria and compared with the gray value-based symmetry measure and two other methods from the literature. Stability is tested against noise, object rotation, and variations of lighting. As a measure for the distinctiveness of FPs, the principal components of FP-centered windows are compared with those of windows at randomly chosen points on a large database of natural images. Finally, usability is evaluated in the context of an object recognition task.
Gunther Heidemann
IEEE Trans. Pattern Anal. Mach. Intell.1
2003 Semi-automatic acquisition and labelling of image data using SOMs
Gunther Heidemann, Axel Saalbach, Helge J. Ritter
ESANN1
2003 Recognition of Gestural Object Reference with Auditory Feedback
Ingo Bax, Holger Bekel, Gunther Heidemann
ICANN3
2003 Integrating Context-Free and Context-Dependent Attentional Mechanisms for Gestural Object Reference
Gunther Heidemann, Robert Rae, Holger Bekel, Ingo Bax, Helge J. Ritter
ICVS1
2002 Combining gestural and contact information for visual guidance of multi-finger grasps
Gunther Heidemann, Helge J. Ritter
ESANN1
2002 Parametrized SOMs for Object Recognition and Pose Estimation
Axel Saalbach, Gunther Heidemann, Helge J. Ritter
ICANN2
2001 Visual Checking of Grasping Positions of a Three-Fingered Robot Hand
Gunther Heidemann, Helge J. Ritter
ICANN1
2001 Guiding attention for grasping tasks by gestural instruction: the GRAVIS-robot architecture
abstract
A major goal for the realization of a new generation of intelligent robots is the capability of instructing work tasks by interactive demonstration. To make such a process efficient and convenient for the human user requires that both the robot and the user can establish and maintain a common focus of attention. We describe a hybrid architecture that combines neural networks and finite stale machines into a flexible framework for controlling the behaviour of a vision based robot called GRAVIS-robot (Gestural Recognition Active Vision System robot). It consists of a binocular camera head, a 6 DOF robot arm and a 9 DOF multifingered hand. We focus primarily on nonverbal communication based on gestural commands of a human instructor which will at a later stage be complemented by spoken instructions.
Jochen J. Steil, Gunther Heidemann, Ján Jockusch, Robert Rae, Nils Jungclaus, Helge J. Ritter
IROS2
2001 Efficient Vector Quantization Using the WTA-Rule with Activity Equalization
Gunther Heidemann, Helge J. Ritter
Neural Process. Lett.1
2000 A System for Various Visual Classification Tasks Based on Neural Networks
abstract
A three stage recognition architecture that can be trained to different recognition or segmentation tasks is presented. It consists of an adaptive feature extraction based on vector quantization and local PCA. The features are classified by neural expert networks. It is shown that the system can be applied to object classification, segmentation of partially occluded objects and classification of object parts without modifications in the architecture.
Gunther Heidemann, Dirk Lücke, Helge J. Ritter
ICPR1
2000 Segmentation of Partially Occluded Objects by Local Classification
abstract
An algorithm for supervised learning the segmentation of partially occluded objects is presented. It is based on the classification of object windows which are small compared to the object size but large enough to evaluate structural object features as well as colour. From the input windows, features are extracted by local principal component analysis and subsequently classified by a neural network of the local linear map type. The performance is checked on images of objects with partial occlusion which were artificially generated from the Columbia Object Image Library.
Gunther Heidemann, Dirk Lücke, Helge J. Ritter
IJCNN (1)1
1999 Artificial neural networks for automated quality control of textile seams
Claus Bahlmann, Gunther Heidemann, Helge J. Ritter
Pattern Recognit.2
1999 A multi-directional multiple-path recognition scheme for complex objects applied to the domain of a wooden toy kit
Elke Braun, Gunther Heidemann, Helge J. Ritter, Gerhard Sagerer
Pattern Recognit. Lett.2
1996 A Hybrid Object Recognition Architecture
Gunther Heidemann, Franz Kummert, Helge J. Ritter, Gerhard Sagerer
ICANN1
1996 A neural 3-D object recognition architecture using optimized Gabor filters
abstract
We present an object recognition architecture based on feature extraction by Gabor filter kernels and feature classification by an artificial neural network. The parameters of the Gabor filters are optimized to the specific problem by minimizing an energy function. Such Gabor filters extract features that can be more easily classified by the neural network. Moreover, the feature space is low-dimensional so feature extraction does not require much computational effort. The object recognition system is implemented on a Datacube and works in real-time.
Gunther Heidemann, Helge J. Ritter
ICPR1