EDBT 2026 Demo / reviewers in the wild / expert
Heiko Neumann
dblp:n/HeikoNeumann
· DBLP profile ↗
71ranked-venue papers
6as first author
10since 2021 · last 2025
0000-0001-7687-5792ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 6 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 3 since 2021Systems, architecture and hardware · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
8 papers |
Video understanding and tracking · 31% 3D vision · 30% Robot manipulation · 20% | |
| Computer graphics and multimedia
4 papers |
Image and video processing · 76% Geometric modeling and processing · 24% |
Topics — the 21 heaviest of 26, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › 3D vision › local feature descriptor › descriptor learning
dense object descriptors |
1.3 | 2 | 2024 | Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images · ICRA 2024 Efficient and Robust Training of Dense Object Nets for Multi-Object Robot Manipulation · ICRA 2022 |
Robotics › Robot manipulation
grasping |
1.3 | 2 | 2024 | Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images · ICRA 2024 Efficient and Robust Training of Dense Object Nets for Multi-Object Robot Manipulation · ICRA 2022 |
Computer vision › Video understanding and tracking › video object segmentation
online video object segmentation |
0.9 | 1 | 2025 | A Filtering Framework for Semi-online Referring Video Object Segmentation · ACM Multimedia 2025 |
Computer vision › Segmentation and scene understanding
referring image segmentation |
0.9 | 1 | 2025 | A Filtering Framework for Semi-online Referring Video Object Segmentation · ACM Multimedia 2025 |
Computer vision › Video understanding and tracking
video object segmentation |
0.9 | 1 | 2025 | A Filtering Framework for Semi-online Referring Video Object Segmentation · ACM Multimedia 2025 |
Computer vision › 3D vision › feature matching › local feature matching
keypoint matching |
0.8 | 1 | 2024 | Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images · ICRA 2024 |
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › self-supervised visual representation learning
self-supervised correspondence learning |
0.8 | 1 | 2024 | Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB Images · ICRA 2024 |
Robotics › Robot manipulation › object manipulation
multi-object manipulation |
0.6 | 1 | 2022 | Efficient and Robust Training of Dense Object Nets for Multi-Object Robot Manipulation · ICRA 2022 |
Computer vision › 3D vision › 3d generation
3d human generation |
0.4 | 1 | 2020 | Generating 3D People in Scenes Without People · CVPR 2020 |
Computer vision › Video understanding and tracking › dynamic scene analysis › video scene understanding › human-centric scene understanding
human-scene interaction |
0.4 | 1 | 2020 | Generating 3D People in Scenes Without People · CVPR 2020 |
Computer vision › Video understanding and tracking
action segmentation |
0.4 | 1 | 2019 | Local Temporal Bilinear Pooling for Fine-Grained Action Parsing · CVPR 2019 |
Computer vision › 3D vision
object pose estimation |
0.2 | 1 | 2022 | Efficient and Robust Training of Dense Object Nets for Multi-Object Robot Manipulation · ICRA 2022 |
Machine learning › Generative modeling › variational autoencoder
conditional variational autoencoder |
0.1 | 1 | 2020 | Generating 3D People in Scenes Without People · CVPR 2020 |
Image and video processing
motion analysis |
0.1 | 1 | 2007 | Disambiguating Visual Motion by Form-Motion Interaction - a Computational Model · Int. J. Comput. Vis. 2007 |
Image and video processing
motion estimation |
0.1 | 1 | 2007 | A Fast Biologically Inspired Algorithm for Recurrent Motion Estimation · IEEE Trans. Pattern Anal. Mach. Intell. 2007 |
Geometric modeling and processing
tensor voting |
0.0 | 1 | 2004 | Are Iterations and Curvature Useful for Tensor Voting? · ECCV (3) 2004 |
Machine learning › Deep learning architectures and training › attention mechanism
visual attention |
0.0 | 1 | 1999 | Space-Variant Dynamic Neural Fields for Visual Attention · CVPR 1999 |
Bioinformatics and computational biology
computational neuroscience |
0.0 | 1 | 1997 | Detection of First and Second Order Motion · NIPS 1997 |
Bioinformatics and computational biology › computational neuroscience › sensory processing
motion perception |
0.0 | 1 | 1997 | Detection of First and Second Order Motion · NIPS 1997 |
Computational geometry › differential geometry
curvature |
0.0 | 1 | 2004 | Are Iterations and Curvature Useful for Tensor Voting? · ECCV (3) 2004 |
Robotics › Robot navigation and mapping › active vision
gaze control |
0.0 | 1 | 1999 | Space-Variant Dynamic Neural Fields for Visual Attention · CVPR 1999 |
Methods — techniques the papers use, named apart from their topics
temporal consistency · 0.9stochastic optimization · 0.9filtering · 0.9self-supervision · 0.8cycle consistency · 0.8dense object nets · 0.6data augmentation · 0.6surface-based human model · 0.4scene constraint optimization · 0.4conditional variational autoencoder · 0.4sparse coding · 0.1shunting inhibition · 0.1feedback modulation · 0.1computational model of form-motion interaction · 0.1motion energy analysis · 0.0computational modeling · 0.0ocular stripe maps · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | A Filtering Framework for Semi-online Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) extracts objects from videos based on provided text narrations. Previous approaches typically work on all video frames simultaneously through offline processing, but this is not always possible. Offline processing becomes also ineffective for long videos, as directly cutting the video into short clips to fit memory limits results in the loss of temporal consistency. To make RVOS more applicable to real-world video streams or long video scenarios, we introduce a Filtering Framework for RVOS (FF-RVOS), the first model capable of operating in online, semi-online, and offline modes with just one training session. We redesign RVOS as a stochastic optimization problem and leverage filtering to optimize object states across temporal sequences. Our method enhances temporal consistency for video streams and when a long video is processed in shorter clip sequences due to memory limitations. FF-RVOS demonstrates superior performance compared to previous state-of-the-art methods on public benchmarks with clear improvements, especially when the video is cut into short clips for processing. Our framework can also be embedded into different offline methods to boost temporal consistency. The project page is https://github.com/haliphinx/FF-RVOS. Xiao Hu 0008, Heiko Neumann, Jochen Lang 0001 |
ACM Multimedia | 2 |
| 2025 | A model of thalamo-cortical interaction for incremental binding in mental contour-tracingabstractObject-basd visual attention marks a key process of mammalian perception. By which mechanisms this process is implemented and how it can be interacted with by means of attentional control is not completely understood yet. Incremental binding is a mechanism required in demanding scenarios of object-based attention and is experimentally well investigated. Attention spreads across a representation of the visual object and labels bound elements by constant up-modulation of neural activity. The speed of incremental binding was found to be dependent on the spatial arrangement of distracting elements in the scene and to be scale invariant giving rise to the growth-cone hypothesis. In this work, we propose a neural dynamical model of incremental binding that provides a mechanistic account for these findings. Through simulations, we investigate the model properties and demonstrate how an attentional spreading mechanism tags neurons that participate in the object binding process. They utilize Gestalt properties and eventually show growth-cone characteristics labeling perceptual items by delayed activity enhancement of neuronal firing rates. We discuss the algorithmic process underlying incremental binding and relate it to our model computations. This theoretical investigation encompasses complexity considerations and finds the model to be not only of explanatory value in terms of neurophysiological evidence, but also to be an efficient implementation of incremental binding striving to establish a normative account. By relating the connectivity motifs of the model to neuroanatomical evidence, we suggest thalamo-cortical interactions to be a likely candidate for the flexible and efficient realization suggested by the model. There, pyramidal cells are proposed to serve as the processors of incremental grouping information. Local bottom-up evidence about stimulus features is integrated via basal dendritic sites. It is combined with an apical signal consisting of contextual grouping information which is gated by attentional task-relevance selection mediated via higher-order thalamic representations. Daniel Schmid, Heiko Neumann |
PLoS Comput. Biol. | 2 |
| 2024 | Cycle-Correspondence Loss: Learning Dense View-Invariant Visual Features from Unlabeled and Unordered RGB ImagesabstractRobot manipulation relying on learned object-centric descriptors became popular in recent years. Visual descriptors can easily describe manipulation task objectives, they can be learned efficiently using self-supervision, and they can encode actuated and even non-rigid objects. However, learning robust, view-invariant keypoints in a self-supervised approach requires a meticulous data collection approach involving precise calibration and expert supervision. In this paper we introduce Cycle-Correspondence Loss (CCL) for view-invariant dense descriptor learning, which adopts the concept of cycle-consistency, enabling a simple data collection pipeline and training on unpaired RGB camera views. The key idea is to autonomously detect valid pixel correspondences by attempting to use a prediction over a new image to predict the original pixel in the original image, while scaling error terms based on the estimated confidence. Our evaluation shows that we outperform other self-supervised RGB-only methods, and approach performance of supervised methods, both with respect to keypoint tracking as well as for a robot grasping downstream task. David B. Adrian, Andras Gabor Kupcsik, Markus Spies, Heiko Neumann |
ICRA | 4 |
| 2024 | Temporal Context Enhanced Referring Video Object SegmentationabstractThe goal of Referring Video Object Segmentation is to extract an object from a video clip based on a given expression. While previous methods have utilized the transformer’s multi-modal learning capabilities to aggregate information from different modalities, they have mainly focused on spatial information and paid less attention to temporal information. To enhance the learning of temporal information, we propose TCE-RVOS with a novel frame token fusion (FTF) structure and a novel instance query transformer (IQT). Our technical innovations maximize the potential information gain of videos over single images. Our contributions also include a new classification of two widely used validation datasets for investigation of challenging cases. Our experimental results demonstrate that TCERVOS effectively captures temporal information and outperforms the previous state-of-the-art methods by increasing the J&F score by 4.0 and 1.9 points using ResNet-50 and VSwin-Tiny as the backbone on Ref-Youtube-VOS, respectively, and +2.0 mAP on A2D-Sentences dataset by using VSwin-Tiny backbone. The code is available at https://github.com/haliphinx/TCE-RVOS Xiao Hu 0008, Basavaraj Hampiholi, Heiko Neumann, Jochen Lang 0001 |
WACV | 3 |
| 2024 | Teaching deep networks to see shape: Lessons from a simplified visual worldabstractDeep neural networks have been remarkably successful as models of the primate visual system. One crucial problem is that they fail to account for the strong shape-dependence of primate vision. Whereas humans base their judgements of category membership to a large extent on shape, deep networks rely much more strongly on other features such as color and texture. While this problem has been widely documented, the underlying reasons remain unclear. We design simple, artificial image datasets in which shape, color, and texture features can be used to predict the image class. By training networks from scratch to classify images with single features and feature combinations, we show that some network architectures are unable to learn to use shape features, whereas others are able to use shape in principle but are biased towards the other features. We show that the bias can be explained by the interactions between the weight updates for many images in mini-batch gradient descent. This suggests that different learning algorithms with sparser, more local weight changes are required to make networks more sensitive to shape and improve their capability to describe human vision. Christian Jarvers, Heiko Neumann |
PLoS Comput. Biol. | 2 |
| 2023 | Defining gaze patterns for process model literacy - Exploring visual routines in process models with diverse mappings
Michael Winter 0002, Heiko Neumann, Rüdiger Pryss, Thomas Probst, Manfred Reichert |
Expert Syst. Appl. | 2 |
| 2023 | Corrigendum to "Defining gaze patterns for process model literacy - Exploring visual routines in process models with diverse mappings" [Expert Syst. Appl. 213 (2023) 119217]
Michael Winter 0002, Heiko Neumann, Rüdiger Pryss, Thomas Probst, Manfred Reichert |
Expert Syst. Appl. | 2 |
| 2022 | Efficient and Robust Training of Dense Object Nets for Multi-Object Robot ManipulationabstractWe propose a framework for robust and efficient training of Dense Object Nets (DON) [1] with a focus on industrial multi-object robot manipulation scenarios. DON is a popular approach to obtain dense, view-invariant object descriptors, which can be used for a multitude of downstream tasks in robot manipulation, such as, pose estimation, state representation for control, etc. However, the original work [1] focused training on singulated objects, with limited results on instance-specific, multi-object applications. Additionally, a complex data collection pipeline, including 3D reconstruction and mask annotation of each object, is required for training. In this paper, we further improve the efficacy of DON with a simplified data collection and training regime, that consistently yields higher precision and enables robust tracking of keypoints with less data requirements. In particular, we focus on training with multi-object data instead of singulated objects, combined with a well-chosen augmentation scheme. We additionally propose an alternative loss formulation to the original pixel wise formulation that offers better results and is less sensitive to hyperparameters. Finally, we demonstrate the robustness and accuracy of our proposed framework on a real-world robotic grasping task. David B. Adrian, Andras Gabor Kupcsik, Markus Spies, Heiko Neumann |
ICRA | 4 |
| 2021 | How Healthcare Professionals Comprehend Process Models - An Empirical Eye Tracking AnalysisabstractDigitization is advancing rapidly in many prevalently analogue domains such as healthcare. For the latter domain, the synergies with modern information technologies (IT) have become an integral part regarding communication and collaboration. For this reason, a comprehensible language is of importance in order to allow a frictionless exchange of information between domain experts. The Business Process Model and Notation (BPMN) 2.0 represents a promising notation that may be applied as lingua franca. Although the BPMN 2.0 is widespread applied by experts in business and industry, little experience exists how BPMN 2.0 is adopted in healthcare. In order to assess how BPMN 2.0 is deployed in healthcare, we conducted a preliminary eye tracking study, in which n=16 professionals from healthcare comprehended a particular BPMN 2.0 process model. The results indicate that BPMN 2.0 might be a candidate for a lingua franca to foster the comprehensible exchange of information as well as collaboration between healthcare and IT. Michael Winter 0002, Cynthia Bredemeyer, Manfred Reichert, Heiko Neumann, Thomas Probst, Rüdiger Pryss |
CBMS | 4 |
| 2021 | Multi-Modal Pain Intensity Recognition Based on the SenseEmotion DatabaseabstractThe subjective nature of pain makes it a very challenging phenomenon to assess. Most of the current pain assessment approaches rely on an individual’s ability to recognise and report an observed pain episode. However, pain perception and expression are affected by numerous factors ranging from personality traits to physical and psychological health state. Hence, several approaches have been proposed for the automatic recognition of pain intensity, based on measurable physiological and audiovisual parameters. In the current paper, an assessment of several fusion architectures for the development of a multi-modal pain intensity classification system is performed. The contribution of the presented work is two-fold: (1) 3 distinctive modalities consisting of audio, video and physiological channels are assessed and combined for the classification of several levels of pain elicitation. (2) An extensive assessment of several fusion strategies is carried out in order to design a classification architecture that improves the performance of the pain recognition system. The assessment is based on theSenseEmotion Databaseand experimental validation demonstrates the relevance of the multi-modal classification approach, which achieves classification rates of respectively$83.39\%$,$59.53\%$and$43.89\%$in a 2-class, 3-class and 4-class pain intensity classification task. Patrick Thiam, Viktor Kessler, Mohammadreza Amirian, Peter Bellmann, Georg Layher, Yan Zhang 0054, Maria Velana, Sascha Gruss, Steffen Walter 0001, Harald C. Traue, Daniel Schork, Jonghwa Kim 0001, Elisabeth André, Heiko Neumann, Friedhelm Schwenker |
IEEE Trans. Affect. Comput. | 14 |
| 2020 | Depthwise Separable Temporal Convolutional Network for Action SegmentationabstractFine-grained temporal action segmentation in long, untrimmed RGB videos is a key topic in visual human-machine interaction. Recent temporal convolution based approaches either use encoder-decoder(ED) architecture or dilations with doubling factor in consecutive convolution layers to segment actions in videos. However ED networks operate on low temporal resolution and the dilations in successive layers cause gridding artifacts problem. We propose depthwise separable temporal convolution network (DS-TCN) that operates on full temporal resolution and with reduced gridding effects. The basic component of DS-TCN is residual depthwise dilated block (RDDB). We explore the trade-off between large kernels and small dilation rates using RDDB. We show that our DS-TCN is capable of capturing long-term dependencies as well as local temporal cues efficiently. Our evaluation on three benchmark datasets, GTEA, 50Salads, and Breakfast demonstrates that DS-TCN outperforms the existing ED-TCN and dilation based TCN baselines even with comparatively fewer parameters. Basavaraj Hampiholi, Christian Jarvers, Wolfgang Mader 0003, Heiko Neumann |
3DV | 4 |
| 2020 | Generating 3D People in Scenes Without PeopleabstractWe present a fully automatic system that takes a 3D scene and generates plausible 3D human bodies that are posed naturally in that 3D scene. Given a 3D scene without people, humans can easily imagine how people could interact with the scene and the objects in it. However, this is a challenging task for a computer as solving it requires that (1) the generated human bodies to be semantically plausible within the 3D environment (e.g. people sitting on the sofa or cooking near the stove), and (2) the generated human-scene interaction to be physically feasible such that the human body and scene do not interpenetrate while, at the same time, body-scene contact supports physical interactions. To that end, we make use of the surface-based 3D human model SMPL-X. We first train a conditional variational autoencoder to predict semantically plausible 3D human poses conditioned on latent scene representations, then we further refine the generated 3D bodies using scene constraints to enforce feasible physical interaction. We show that our approach is able to synthesize realistic and expressive 3D human bodies that naturally interact with 3D environment. We perform extensive experiments demonstrating that our generative framework compares favorably with existing methods, both qualitatively and quantitatively. We believe that our scene-conditioned 3D human generation pipeline will be useful for numerous applications; e.g. to generate training data for human pose estimation, in video games and in VR/AR. Our project page for data and code can be seen at: {https://vlg.inf.ethz.ch/projects/PSI/}. Yan Zhang 0054, Mohamed Hassan 0003, Heiko Neumann, Michael J. Black, Siyu Tang 0001 |
CVPR | 3 |
| 2020 | Classifier-Guided Visual Correction of Noisy Labels for Image Classification TasksabstractAbstract Training data plays an essential role in modern applications of machine learning. However, gathering labeled training data is time‐consuming. Therefore, labeling is often outsourced to less experienced users, or completely automated. This can introduce errors, which compromise valuable training data, and lead to suboptimal training results. We thus propose a novel approach that uses the power of pretrained classifiers to visually guide users to noisy labels, and let them interactively check error candidates, to iteratively improve the training data set. To systematically investigate training data, we propose a categorization of labeling errors into three different types, based on an analysis of potential pitfalls in label acquisition processes. For each of these types, we present approaches to detect, reason about, and resolve error candidates, as we propose measures and visual guidance techniques to support machine learning users. Our approach has been used to spot errors in well‐known machine learning benchmark data sets, and we tested its usability during a user evaluation. While initially developed for images, the techniques presented in this paper are independent of the classification algorithm, and can also be extended to many other types of training data. Alex Bäuerle, Heiko Neumann, Timo Ropinski |
Comput. Graph. Forum | 2 |
| 2020 | Computational principles of neural adaptation for binaural signal integrationabstractAdaptation to statistics of sensory inputs is an essential ability of neural systems and extends their effective operational range. Having a broad operational range facilitates to react to sensory inputs of different granularities, thus is a crucial factor for survival. The computation of auditory cues for spatial localization of sound sources, particularly the interaural level difference (ILD), has long been considered as a static process. Novel findings suggest that this process of ipsi- and contra-lateral signal integration is highly adaptive and depends strongly on recent stimulus statistics. Here, adaptation aids the encoding of auditory perceptual space of various granularities. To investigate the mechanism of auditory adaptation in binaural signal integration in detail, we developed a neural model architecture for simulating functions of lateral superior olive (LSO) and medial nucleus of the trapezoid body (MNTB) composed of single compartment conductance-based neurons. Neurons in the MNTB serve as an intermediate relay population. Their signal is integrated by the LSO population on a circuit level to represent excitatory and inhibitory interactions of input signals. The circuit incorporates an adaptation mechanism operating at the synaptic level based on local inhibitory feedback signals. The model's predictive power is demonstrated in various simulations replicating physiological data. Incorporating the innovative adaptation mechanism facilitates a shift in neural responses towards the most effective stimulus range based on recent stimulus history. The model demonstrates that a single LSO neuron quickly adapts to these stimulus statistics and, thus, can encode an extended range of ILDs in the ipsilateral hemisphere. Most significantly, we provide a unique measurement of the adaptation efficacy of LSO neurons. Prerequisite of normal function is an accurate interaction of inhibitory and excitatory signals, a precise encoding of time and a well-tuned local feedback circuit. We suggest that the mechanisms of temporal competitive-cooperative interaction and the local feedback mechanism jointly sensitize the circuit to enable a response shift towards contra-lateral and ipsi-lateral stimuli, respectively. Timo Oess, Marc O. Ernst, Heiko Neumann |
PLoS Comput. Biol. | 3 |
| 2019 | Local Temporal Bilinear Pooling for Fine-Grained Action ParsingabstractFine-grained temporal action parsing is important in many applications, such as daily activity understanding, human motion analysis, surgical robotics and others requiring subtle and precise operations over a long-term period. In this paper we propose a novel bilinear pooling operation, which is used in intermediate layers of a temporal convolutional encoder-decoder net. In contrast to previous work, our proposed bilinear pooling is learnable and hence can capture more complex local statistics than the conventional counterpart. In addition, we introduce exact lower-dimension representations of our bilinear forms, so that the dimensionality is reduced without suffering from information loss nor requiring extra computation. We perform extensive experiments to quantitatively analyze our model and show the superior performances to other state-of-the-art pooling work on various datasets. Yan Zhang 0054, Siyu Tang 0001, Krikamol Muandet, Christian Jarvers, Heiko Neumann |
CVPR | 5 |
| 2019 | Temporal Learning of Dynamics in Complex Neuron Models using BackpropagationabstractOne of the major challenges of computational cognitive neuroscience is to apply models of neural information processing to complex tasks. Hierarchical learning architectures like deep convolutional networks can be trained to solve tasks efficiently, but utilize simple mechanisms of activity integration and output generation. On the other hand, biologically plausible models of activation dynamics incorporate detailed mechanisms of changing membrane potentials and axonal firing properties. Making such elaborate models trainable requires learning of internal model parameters. Here, we propose to apply supervised learning and train a model of canonical cortical circuits via backpropagation through time. We train the model to settle to target equilibrium values, to generate oscillations, and to solve a contour completion task. Christian Jarvers, Daniel Schmid, Heiko Neumann |
IJCNN | 3 |
| 2019 | Motion Integration and Disambiguation by Spiking V1-MT-MSTl Feedforward-Feedback InteractionabstractMotion detection registers items within restricted regions in the visual field. Early stages of cortical processing of motion advance this estimate by integrating spatio-temporal input responses in area V1 to build feature representations of direction and speed in area MT of primate cortex. The neural mechanisms underlying such processes are not yet fully understood. We propose a neural model of hierarchically organized areas V1, MT, and MSTl, with feedforward and feedback connections. Each area serves a distinct purpose and is formally represented by layers of model cortical columns composed of excitatory and inhibitory spiking neurons with conductance-based activation dynamics. Recurrent connections enhance activations by modulatory interaction and divisive normalization. MT population activities allow to estimate motion direction and speed which we show for various stimuli. The importance of the feedback connections for disambiguation is demonstrated in simulated lesion studies. Maximilian P. R. Löhr, Daniel Schmid, Heiko Neumann |
IJCNN | 3 |
| 2018 | Human Motion Parsing by Hierarchical Dynamic Clustering
Yan Zhang 0054, Siyu Tang 0001, Heiko Neumann |
BMVC | 4 |
| 2018 | Contrast Detection in Event-Streams from Dynamic Vision Sensors with Fixational Eye MovementsabstractThe purpose of fixational eye movements of the human eye has long been - and still is - an intensively discussed topic in the research community. Such micro-movements seem to prevent fading, may sharpen the visual image by increasing high-frequency signal components and help to counteract the eye drifting away from its target. We investigate the impact of these movements for event-based visual sensing with a DVS128 camera. Events are only generated whenever the luminance at a pixel changes and, therefore, static form cannot be resolved in a stationary camera. We present a specially designed mirror system that creates virtual fixational eye movements for an event-based camera. In addition, we investigate the shape of the Fourier spectrum of random miniature motions of the sensory recordings for stationary as well as moving contrast features that give rise to ON/OFF event responses. From this we draw conlusions concerning the optimal shape of filters to detect local motion. Maximilian P. R. Löhr, Heiko Neumann |
ISCAS | 2 |
| 2018 | Facial point localization via neural networks in a cascade regression framework
Anwar Saeed, Ayoub Al-Hamadi, Heiko Neumann |
Multim. Tools Appl. | 3 |
| 2017 | Visual Confusion Recognition in Movement Patterns from Walking Path and Motion Energy
Yan Zhang 0054, Georg Layher, Steffen Walter 0001, Viktor Kessler, Heiko Neumann |
ICOST | 5 |
| 2016 | Bio-inspired computer vision: Towards a synergistic approach of artificial and biological visionabstractStudies in biological vision have always been a great source of inspiration for design of computer vision algorithms . In the past, several successful methods were designed with varying degrees of correspondence with biological vision studies, ranging from purely functional inspiration to methods that utilise models that were primarily developed for explaining biological observations. Even though it seems well recognised that computational models of biological vision can help in design of computer vision algorithms, it is a non-trivial exercise for a computer vision researcher to mine relevant information from biological vision literature as very few studies in biology are organised at a task level. In this paper we aim to bridge this gap by providing a computer vision task centric presentation of models primarily originating in biological vision studies. Not only do we revisit some of the main features of biological vision and discuss the foundations of existing computational studies modelling biological vision, but also we consider three classical computer vision tasks from a biological perspective: image sensing, segmentation and optical flow. Using this task-centric approach, we discuss well-known biological functional principles and compare them with approaches taken by computer vision. Based on this comparative analysis of computer and biological vision, we present some recent models in biological vision and highlight a few models that we think are promising for future investigations in computer vision. To this extent, this paper provides new insights and a starting point for investigators interested in the design of biology-based computer vision algorithms and pave a way for much needed interaction between the two communities leading to the development of synergistic models of artificial and biological vision. N. V. Kartheek Medathati, Heiko Neumann, Guillaume S. Masson, Pierre Kornprobst |
Comput. Vis. Image Underst. | 2 |
| 2015 | Towards the Separation of Rigid and Non-rigid Motions for Facial Expression AnalysisabstractIn intelligent environments, computer systems not solely serve as passive input devices waiting for user interaction but actively analyze their environment and adapt their behaviour according to changes in environmental parameters. One essential ability to achieve this goal is to analyze the mood, emotions and dispositions a user experiences while interacting with such intelligent systems. Features allowing to infer such parameters can be extracted from auditive, as well as visual sensory input streams. For the visual feature domain, in particular facial expressions are known to contain rich information about a user's emotional state and can be detected by using either static and/or dynamic image features. During interaction facial expressions are rarely performed in isolation, but most of the time co-occur with movements of the head. Thus, optical flow based facial features are often compromised by additional motions. Parts of the optical flow may be caused by rigid head motions, while other parts reflect deformations resulting from facial expressivity (non-rigid motions). In this work, we propose the first steps towards an optical flow based separation of rigid head motions from non-rigid motions caused by facial expressions. We suggest that after their separation, both, head movements and facial expressions can be used as a basis for the recognition of a user's emotions and dispositions and thus allow a technical system to effectively adapt to the user's state. Georg Layher, Stephan Tschechne, Robert Niese, Ayoub Al-Hamadi, Heiko Neumann |
Intelligent Environments | 5 |
| 2015 | Reinforcement Learning of Linking and Tracing Contours in Recurrent Neural NetworksabstractThe processing of a visual stimulus can be subdivided into a number of stages. Upon stimulus presentation there is an early phase of feedforward processing where the visual information is propagated from lower to higher visual areas for the extraction of basic and complex stimulus features. This is followed by a later phase where horizontal connections within areas and feedback connections from higher areas back to lower areas come into play. In this later phase, image elements that are behaviorally relevant are grouped by Gestalt grouping rules and are labeled in the cortex with enhanced neuronal activity (object-based attention in psychology). Recent neurophysiological studies revealed that reward-based learning influences these recurrent grouping processes, but it is not well understood how rewards train recurrent circuits for perceptual organization. This paper examines the mechanisms for reward-based learning of new grouping rules. We derive a learning rule that can explain how rewards influence the information flow through feedforward, horizontal and feedback connections. We illustrate the efficiency with two tasks that have been used to study the neuronal correlates of perceptual organization in early visual cortex. The first task is called contour-integration and demands the integration of collinear contour elements into an elongated curve. We show how reward-based learning causes an enhancement of the representation of the to-be-grouped elements at early levels of a recurrent neural network, just as is observed in the visual cortex of monkeys. The second task is curve-tracing where the aim is to determine the endpoint of an elongated curve composed of connected image elements. If trained with the new learning rule, neural networks learn to propagate enhanced activity over the curve, in accordance with neurophysiological data. We close the paper with a number of model predictions that can be tested in future neurophysiological and computational studies. Tobias Brosch, Heiko Neumann, Pieter R. Roelfsema |
PLoS Comput. Biol. | 2 |
| 2014 | Modeling Simultanagnosia
Anna Belardinelli, Johannes Kurz, Esther Kutter, Heiko Neumann, Hans-Otto Karnath, Martin V. Butz |
CogSci | 4 |
| 2014 | Modeling Perspective-Taking by Correlating Visual and Proprioceptive Dynamics
Fabian Schrodt, Georg Layher, Heiko Neumann, Martin V. Butz |
CogSci | 3 |
| 2014 | Computing with a Canonical Neural Circuits Model with Pool Normalization and Modulating FeedbackabstractEvidence suggests that the brain uses an operational set of canonical computations like normalization, input filtering, and response gain enhancement via reentrant feedback. Here, we propose a three-stage columnar architecture of cascaded model neurons to describe a core circuit combining signal pathways of feedforward and feedback processing and the inhibitory pooling of neurons to normalize the activity. We present an analytical investigation of such a circuit by first reducing its detail through the lumping of initial feedforward response filtering and reentrant modulating signal amplification. The resulting excitatory-inhibitory pair of neurons is analyzed in a 2D phase-space. The inhibitory pool activation is treated as a separate mechanism exhibiting different effects. We analyze subtractive as well as divisive (shunting) interaction to implement center-surround mechanisms that include normalization effects in the characteristics of real neurons. Different variants of a core model architecture are derived and analyzed--in particular, individual excitatory neurons (without pool inhibition), the interaction with an inhibitory subtractive or divisive (i.e., shunting) pool, and the dynamics of recurrent self-excitation combined with divisive inhibition. The stability and existence properties of these model instances are characterized, which serve as guidelines to adjust these properties through proper model parameterization. The significance of the derived results is demonstrated by theoretical predictions of response behaviors in the case of multiple interacting hypercolumns in a single and in multiple feature dimensions. In numerical simulations, we confirm these predictions and provide some explanations for different neural computational properties. Among those, we consider orientation contrast-dependent response behavior, different forms of attentional modulation, contrast element grouping, and the dynamic adaptation of the silent surround in extraclassical receptive field configurations, using only slight variations of the same core reference model. Tobias Brosch, Heiko Neumann |
Neural Comput. | 2 |
| 2014 | Interaction of feedforward and feedback streams in visual cortex in a firing-rate model of columnar computations
Tobias Brosch, Heiko Neumann |
Neural Networks | 2 |
| 2013 | Learning Representations of Animated Motion Sequences - A Neural Model
Georg Layher, Martin A. Giese, Heiko Neumann |
CogSci | 3 |
| 2013 | Attention-Gated Reinforcement Learning in Neural Networks - A Unified View
Tobias Brosch, Friedhelm Schwenker, Heiko Neumann |
ICANN | 3 |
| 2013 | A Biologically Inspired Model for the Detection of External and Internal Head Motions
Stephan Tschechne, Georg Layher, Heiko Neumann |
ICANN | 3 |
| 2013 | A Bio-Inspired, Computational Model Suggests Velocity Gradients of Optic Flow Locally Encode Ordinal Depth at Surface Borders and Globally They Encode Self-MotionabstractVisual navigation requires the estimation of self-motion as well as the segmentation of objects from the background. We suggest a definition of local velocity gradients to compute types of self-motion, segment objects, and compute local properties of optical flow fields, such as divergence, curl, and shear. Such velocity gradients are computed as velocity differences measured locally tangent and normal to the direction of flow. Then these differences are rotated according to the local direction of flow to achieve independence of that direction. We propose a bio-inspired model for the computation of these velocity gradients for video sequences. Simulation results show that local gradients encode ordinal surface depth, assuming self-motion in a rigid scene or object motions in a nonrigid scene. For translational self-motion velocity, gradients can be used to distinguish between static and moving objects. The information about ordinal surface depth and self-motion can help steering control for visual navigation. Florian Raudies, Stefan Ringbauer, Heiko Neumann |
Neural Comput. | 3 |
| 2012 | Learning Representations for Animated Motion Sequence and Implied Motion Recognition
Georg Layher, Martin A. Giese, Heiko Neumann |
ICANN (1) | 3 |
| 2012 | The Brain's Sequential Parallelism: Perceptual Decision-Making and Early Sensory Responses
Tobias Brosch, Heiko Neumann |
ICONIP (2) | 2 |
| 2012 | The Combination of HMAX and HOGs in an Attention Guided Framework for Object Localization
Tobias Brosch, Heiko Neumann |
ICPRAM (2) | 2 |
| 2012 | An audiovisual political speech analysis incorporating eye-tracking and perception data
Stefan Scherer, Georg Layher, John Kane 0002, Heiko Neumann, Nick Campbell 0001 |
LREC | 4 |
| 2012 | A review and evaluation of methods estimating ego-motion
Florian Raudies, Heiko Neumann |
Comput. Vis. Image Underst. | 2 |
| 2012 | Multiscale Binarization of Gene Expression Data for Reconstructing Boolean NetworksabstractNetwork inference algorithms can assist life scientists in unraveling gene-regulatory systems on a molecular level. In recent years, great attention has been drawn to the reconstruction of Boolean networks from time series. These need to be binarized, as such networks model genes as binary variables (either “expressed” or “not expressed”). Common binarization methods often cluster measurements or separate them according to statistical or information theoretic characteristics and may require many data points to determine a robust threshold. Yet, time series measurements frequently comprise only a small number of samples. To overcome this limitation, we propose a binarization that incorporates measurements at multiple resolutions. We introduce two such binarization approaches which determine thresholds based on limited numbers of samples and additionally provide a measure of threshold validity. Thus, network reconstruction and further analysis can be restricted to genes with meaningful thresholds. This reduces the complexity of network inference. The performance of our binarization algorithms was evaluated in network reconstruction experiments using artificial data as well as real-world yeast expression time series. The new approaches yield considerably improved correct network identification rates compared to other binarization techniques by effectively reducing the amount of candidate networks. Martin Hopfensitz, Christoph Müssel, Christian Wawra, Markus Maucher, Michael Kühl, Heiko Neumann, Hans A. Kestler |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2011 | Multiple Classifier Systems for the Classification of Audio-Visual Emotional States
Michael Glodek, Stephan Tschechne, Georg Layher, Martin Schels, Tobias Brosch, Stefan Scherer, Markus Kächele, Miriam Schmidt, Heiko Neumann, Günther Palm, Friedhelm Schwenker |
ACII (2) | 9 |
| 2011 | A Model of Motion Transparency Processing with Local Center-Surround Interactions and FeedbackabstractMotion transparency occurs when multiple coherent motions are perceived in one spatial location. Imagine, for instance, looking out of the window of a bus on a bright day, where the world outside the window is passing by and movements of passengers inside the bus are reflected in the window. The overlay of both motions at the window leads to motion transparency, which is challenging to process. Noisy and ambiguous motion signals can be reduced using a competition mechanism for all encoded motions in one spatial location. Such a competition, however, leads to the suppression of multiple peak responses that encode different motions, as only the strongest response tends to survive. As a solution, we suggest a local center-surround competition for population-encoded motion directions and speeds. Similar motions are supported, and dissimilar ones are separated, by representing them as multiple activations, which occurs in the case of motion transparency. Psychophysical findings, such as motion attraction and repulsion for motion transparency displays, can be explained by this local competition. Besides this local competition mechanism, we show that feedback signals improve the processing of motion transparency. A discrimination task for transparent versus opaque motion is simulated, where motion transparency is generated by superimposing large field motion patterns of either varying size or varying coherence of motion. The model's perceptual thresholds with and without feedback are calculated. We demonstrate that initially weak peak responses can be enhanced and stabilized through modulatory feedback signals from higher stages of processing. Florian Raudies, Ennio Mingolla, Heiko Neumann |
Neural Comput. | 3 |
| 2010 | A neural model of the temporal dynamics of figure-ground segregation in motion perception
Florian Raudies, Heiko Neumann |
Neural Networks | 2 |
| 2008 | The PIT Corpus of German Multi-Party Dialogues
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Harald C. Traue, Ulrich Weidenbacher |
LREC | 4 |
| 2007 | Integration of Multiple Temporal and Spatial Scales for Robust Optic Flow Estimation in a Biologically Inspired Algorithm
Cornelia Beck, Thomas Gottbehuet, Heiko Neumann |
CAIP | 3 |
| 2007 | Neural Mechanisms for Mid-Level Optical Flow Pattern Detection
Stefan Ringbauer, Pierre Bayerl, Heiko Neumann |
ICANN (2) | 3 |
| 2007 | Disambiguating Visual Motion by Form-Motion Interaction - a Computational Model
Pierre Bayerl, Heiko Neumann |
Int. J. Comput. Vis. | 2 |
| 2007 | A Fast Biologically Inspired Algorithm for Recurrent Motion EstimationabstractWe have previously developed a neurodynamical model of motion segregation in cortical visual area V1 and MT of the dorsal stream. The model explains how motion ambiguities caused by the motion aperture problem can be solved for coherently moving objects of arbitrary size by means of cortical mechanisms. The major bottleneck in the development of a reliable biologically inspired technical system with real-time motion analysis capabilities based on this neural model is the amount of memory necessary for the representation of neural activation in velocity space. We propose a sparse coding framework for neural motion activity patterns and suggest a means by which initial activities are detected efficiently. We realize neural mechanisms such as shunting inhibition and feedback modulation in the sparse framework to implement an efficient algorithmic version of our neural model of cortical motion segregation. We demonstrate that the algorithm behaves similarly to the original neural model and is able to extract image motion from real world image sequences. Our investigation transfers a neuroscience model of cortical motion computation to achieve technologically demanding constraints such as real-time performance and hardware implementation. In addition, the proposed biologically inspired algorithm provides a tool for modeling investigations to achieve acceptable simulation time. Pierre Bayerl, Heiko Neumann |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2007 | Iterated tensor voting and curvature improvement
Sylvain Fischer, Pierre Bayerl, Heiko Neumann, Rafael Redondo, Gabriel Cristóbal |
Signal Process. | 3 |
| 2006 | Wizard-of-Oz Data Collection for Perception and Interaction in Multi-User Environments
Petra-Maria Strauß, Holger Hoffmann, Wolfgang Minker, Heiko Neumann, Günther Palm, Stefan Scherer, Friedhelm Schwenker, Harald C. Traue, Welf Walter, Ulrich Weidenbacher |
LREC | 4 |
| 2006 | Sketching shiny surfaces: 3D shape extraction and depiction of specular surfacesabstractMany materials including water, plastic, and metal have specular surface characteristics. Specular reflections have commonly been considered a nuisance for the recovery of object shape. However, the way that reflections are distorted across the surface depends crucially on 3D curvature, suggesting that they could, in fact, be a useful source of information. Indeed, observers can have a vivid impression of, 3D shape when an object is perfectly mirrored (i.e., the image contains nothing but specular reflections). This leads to the question what are the underlying mechanisms of our visual system to extract this 3D shape information from a perfectly mirrored object. In this paper we propose a biologically motivated recurrent model for the extraction of visual features relevant for the perception of 3D shape information from images of mirrored objects. We qualitatively and quantitatively analyze the results of computational model simulations and show that bidirectional recurrent information processing leads to better results than pure feedforward processing. Furthermore, we utilize the model output to create a rough nonphotorealistic sketch representation of a mirrored object, which emphasizes image features that are mandatory for 3D shape perception (e.g., occluding contour and regions of high curvature). Moreover, this sketch illustrates that the model generates a representation of object features independent of the surrounding scene reflected in the mirrored object. Ulrich Weidenbacher, Pierre Bayerl, Heiko Neumann, Roland W. Fleming |
ACM Trans. Appl. Percept. | 3 |
| 2005 | Recovering real-world images from single-scale boundaries with a novel filling-in architecture
Matthias S. Keil, Gabriel Cristóbal, Thorsten Hansen, Heiko Neumann |
Neural Networks | 4 |
| 2004 | Are Iterations and Curvature Useful for Tensor Voting?
Sylvain Fischer, Pierre Bayerl, Heiko Neumann, Gabriel Cristóbal, Rafael Redondo |
ECCV (3) | 3 |
| 2004 | Disambiguating Visual Motion Through Contextual Feedback ModulationabstractMotion of an extended boundary can be measured locally by neurons only orthogonal to its orientation (aperture problem) while this ambiguity is resolved for localized image features, such as corners or nonocclusion junctions. The integration of local motion signals sampled along the outline of a moving form reveals the object velocity. We propose a new model of V1-MT feedforward and feedback processing in which localized V1 motion signals are integrated along the feedforward path by model MT cells. Top-down feedback from MT cells in turn emphasizes model V1 motion activities of matching velocity by excitatory modulation and thus realizes an attentional gating mechanism. The model dynamics implement a guided filling-in process to disambiguate motion signals through biased on-center, off-surround competition. Our model makes predictions concerning the time course of cells in area MT and V1 and the disambiguation process of activity patterns in these areas and serves as a means to link physiological mechanisms with perceptual behavior. We further demonstrate that our model also successfully processes natural image sequences. Pierre Bayerl, Heiko Neumann |
Neural Comput. | 2 |
| 2004 | Neural Mechanisms for the Robust Representation of JunctionsabstractJunctions provide important cues in various perceptual tasks, such as the determination of occlusion relationships for figure-ground separation, transparency perception, and object recognition, among others. In computer vision, junctions are used in a number of tasks, like point matching for image tracking or correspondence analysis. We propose a biologically motivated approach to junction representation in which junctions are implicitly characterized by high activity for multiple orientations within a cortical hypercolumn. A local measure of circular variance is suggested to extract junction points from this distributed representation. Initial orientation measurements are often fragmented and noisy. A coherent contour representation can be generated by a model of V1 utilizing mechanisms of collinear long-range integration and recurrent interaction. In the model, local oriented contrast estimates that are consistent within a more global context are enhanced while inconsistent activities are suppressed. In a series of computational experiments, we compare junction detection based on the new recurrent model with a feedforward model of complex cells. We show that localization accuracy and positive correctness in the detection of generic junction configurations such as L- and T-junctions is improved by the recurrent long-range interaction. Further, receiver operating characteristics analysis is used to evaluate the detection performance on both synthetic and camera images, showing the superior performance of the new approach. Overall, we propose that nonlocal interactions implemented by known mechanisms within V1 play an important role in detecting higher-order features such as corners and junctions. Thorsten Hansen, Heiko Neumann |
Neural Comput. | 2 |
| 2004 | A simple cell model with dominating opponent inhibition for robust image processing
Thorsten Hansen, Heiko Neumann |
Neural Networks | 2 |
| 2003 | A neural model for heading detection from optic flow
Frank Seifart, Pierre Bayerl, Heiko Neumann |
ESANN | 3 |
| 2003 | A view-based approach for object recognition from image sequences
A. Zehender, Pierre Bayerl, Heiko Neumann |
ESANN | 3 |
| 2003 | Neural mechanisms for segregation and recovering of intrinsic image featuresabstractWe present a single-scale architecture for both segregation and recovering of intrinsic image features and brightness perception. Specifically, a given intensity (or grey scale) image is first analyzed for texture (here defined as small-scale even symmetric features), surfaces (small-scale odd symmetric features) and gradients (large-scale even and odd symmetric features). In this way the image is segregated. Subsequently, textures, surfaces and gradients are recovered by corresponding neural circuits. The proposed architecture may serve as a generic building block for a variety of early vision tasks such as, for example, denoising, efficient coding, as well as mid-level tasks that build on the results from the preceding processing stages. Matthias S. Keil, Gabriel Cristóbal, Heiko Neumann |
ICIP (1) | 3 |
| 1999 | Space-Variant Dynamic Neural Fields for Visual AttentionabstractIn this paper we propose a new method for the fast application of dynamic neural fields (DNF) by utilizing the data reduction properties of space-variant active vision (SVAV). We apply this method to the control of visual attention. Dynamic neural fields have several advantages which are useful for many robot vision tasks, e.g. navigation or gaze-control. The dynamics of lateral interaction between neural units generates well-localized areas of high neural activation, which can be easily detected and used for behavior selection. The major focus of this paper is to drastically reduce the computational expense for the application of two-dimensional DNF. For that purpose, the dynamics of DNF is transformed into a space-variant field representation, defining a new type of DNF, namely space-variant dynamic neural fields (SVDNF). The effectiveness of the proposed method is demonstrated for our integrated monocular space-variant vision system. This system uses SVAV for real-time fixation control, depth-from motion estimation and SVDNF for the control of visual attention. Ingo Ahrns, Heiko Neumann |
CVPR | 2 |
| 1999 | Recurrent V1-V2 interaction for early visual information processing
Heiko Neumann, Wolfgang Sepp |
ESANN | 1 |
| 1998 | Real-time monocular fixation control using the log-polar transformation and a confidence-based similarity measureabstractA new technique for the fixation task is proposed. The proposed technique is view-based and uses a gradient descent technique in order to maximize an intensity-based similarity measure between the reference image and the currently grabbed image during control time. A key constraint for the development of computational mechanisms is the achievement of real-time capability on low-cost hardware. We utilize the complex-logarithmic mapping of input images into log-polar space. In addition to a high compression rate this also supports better fixation performance using a space-variant image representation of the image. We further suggest a confidence based similarity measure to overcome the problem of distortions and partial occlusions during the movement of the camera which is the main problem of view-based approaches. We demonstrate the practicability of our method in a real world experiment which solves the fixation task at a frame rate of 16 Hz on an Intel P200 architecture without any special image processing hardware. Ingo Ahrns, Heiko Neumann |
ICPR | 2 |
| 1998 | Robot navigation by combining central and peripheral optical flow detection on a space-variant mapabstractBiological and technical autonomous agents have to achieve basic behaviors of navigation and obstacle avoidance. If they are confined to a monocular visual sensor the optical flow field induced by egomotion can be evaluated in the periphery for speed and direction control, whereas the central flow indicates imminent collisions. We show how a complex-logarithmic mapping of the image is especially apt for the combined evaluation of the flow field for these navigational tasks. We implemented a robot simulation tool to test reliability, speed, control and robustness of our scheme. Christian Toepfer, Moritz Wende, Gregory Baratoff, Heiko Neumann |
ICPR | 4 |
| 1998 | A neural architecture of brightness perception: non-linear contrast detection and geometry-driven diffusion
Heiko Neumann, Luiz Pessoa, Ennio Mingolla |
Image Vis. Comput. | 1 |
| 1997 | Detection of First and Second Order Motion
Alexander Grunewald, Heiko Neumann |
NIPS | 2 |
| 1996 | Neural model for visual contrast detection
Enno Littmann, Heiko Neumann, Luiz Pessoa |
ESANN | 2 |
| 1996 | Neural Model of Cortical Dynamics in Resonant Boundary Detection and Grouping
Heiko Neumann, Petra Mössner |
ICANN | 1 |
| 1996 | Nonlinear interaction of ON and OFF data streams for the detection of visual structureabstractVisual stimuli lead to neural activity in the retina that is propagated in separate ON and OFF pathways to the cortex. Most models of biological early vision recombine these activity streams by a linear integration at the simple cell level. Based on empirical as well as theoretical investigations we propose a nonlinear recombination circuit that is selectively responsive to contrast magnitude as well as to the sharpness of luminance transition. Simulations with artificial and camera images show a higher positional selectivity for local contrasts than an equivalent linear device. In a multiscale hierarchy the nonlinear circuit produces a unique maximum response in scale-space where scale directly relates to the width of the luminance transition. In order to investigate the biological relevance of the proposed neural circuit, we measured the model sensitivity to luminance gradient reversal in bar stimuli. Our simulations show strong similarity to simple cell recordings in the feline striate cortex. This result further supports the evidence for nonlinear interaction at the simple cell level. Enno Littmann, Heiko Neumann, Luiz Pessoa |
ICPR | 2 |
| 1996 | Mechanisms of Neural Architecture for Visual Contrast and Brightness Perception
Heiko Neumann |
Neural Networks | 1 |
| 1994 | Local stereoscopic depth estimation
Kai-Oliver Ludwig, Heiko Neumann, Bernd Neumann |
Image Vis. Comput. | 2 |
| 1992 | Local Stereoscopic Depth Estimation Using Ocular Stripe Maps
Kai-Oliver Ludwig, Heiko Neumann, Bernd Neumann |
ECCV | 2 |
| 1992 | Estimating attributes of smooth signal transitions from scale-spaceabstractStep-edge models as they have been used to model local intensity variation, only rarely are justified for the real case of image data. Due to finite apertures, the nature of scene geometry as well as discretization of the image, local intensity variations result in smooth transitions of varying width and local contrast. In order to appropriately deal with the robust detection and localization of image contrast, the authors propose the parametrized ramp transition as local signal model. The scale-space processing scheme for token extraction consists of a cascade of first band-pass filtering the raw data and a subsequent correlation of the result with a scaled first order derivative operator. The robust contrast detection within scale space and the estimation of local signal attributes in closed form is documented. The scheme can be extended to deal with intensity variations of different specificity.> Heiko Neumann, Karsten Ottenberg |
ICPR (3) | 1 |
| 1992 | Finding and describing local structure in discrete two-dimensional computed tomogramsabstractDifferently scaled intensity discontinuities along the gradient direction of scaled oriented contrast edges define significant X-ray CT and MR image structures. Reliable detection and quantitative description of such scaled intensity discontinuities is achieved with a scale-space approach. It results in an explicit token representation which can be understood as a compact symbolic description of significant image structure, namely a set of quantitative attributes per contrast edge. Moreover, a multi-resolution approach is used to reconstruct the image intensity function from the token representation alone using a membrane as prior image model. Current results from the novel entire processing cascade, which combines a quantitative discontinuity description and a membrane-based intensity reconstruction, are presented.> Heiko Neumann, Karsten Ottenberg, H. Siegfried Stiehl |
ICPR (3) | 1 |