Erhardt Barth

dblp:90/4529 · DBLP profile ↗
← Back
47ranked-venue papers
1as first author
9since 2021 · last 2026
0000-0001-8556-2472ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 32 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-authorHuman-computer interaction and ubiquitous computing · 3Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AI-based collimation optimization for X-ray imaging using depth cameras
abstract
Collimation during radiography, which is the process of defining the area to be radiated, is a crucial factor for the protection of the patient and for the diagnostic quality of a radiograph. Moreover, incorrect collimation is one of the main causes for a retake and the associated costs. In this paper we propose a novel collimation optimization approach using depth cameras and deep Neural Networks trained end-to-end. We have acquired two new datasets for this purpose. The first, obtained in a clinical environment, consists of depth images of the lower leg and abdomen and the second, captured in real clinical practice, consists of depth images and corresponding radiographs of thorax examinations. For all depth images, the ideal collimation was labeled by experts either on the depth image or directly on the radiograph. Using this dataset to learn to predict the optimal collimation, we show that it is possible to learn different shapes of collimations and to achieve results that are on par with those obtained by radiographers. Such an AI assistant trained with optimal collimation could reduce the radiation exposure to which the patient is exposed, improve the workflow in radiography, and finally increase the diagnostic quality of radiographs.
Dominik Mairhöfer, Manuel Laufer, Lennart Berkel, Malte Sieren, Arpad Bischof, Erhardt Barth, Jörg Barkhausen, Thomas Martinetz
Neurocomputing6
2025 Investigating the Impact of Imbalanced Medical Data on the Performance of Self-Supervised Learning Approaches
abstract
In clinical practice, a substantial amount of data is generated on a daily basis for diagnostic purposes.Since expensive expert knowledge is required for data annotation in order to use this data for supervised learning, large amounts of data often remain unused.Self-supervised learning methods are well suited for using unlabeled data by pre-training networks to solve pretext tasks.As medical data follow an underlying uneven distribution of occurring diseases, they are inherently imbalanced.This could introduce an unwanted bias during pre-training, ultimately leading to negative consequences that may inhibit the benefits of finetuning.In this work we investigate the impact of the imbalance of 2D and 3D medical datasets used for pre-training, as well as the importance of the type and size of the dataset used for pre-training and the pretext task.Our findings indicate that the size of the dataset used for pre-training has greater impact on the final tasks than its balance.
Manuel Laufer, Felicitas Brokmann, Dominik Mairhöfer, Erhardt Barth, Thomas Martinetz
ESANN4
2024 AI-based Collimation Optimization for X-Ray Imaging using Time-of-Flight Cameras
abstract
Collimation during radiography, which is the process of defining the area to be radiated, is a crucial factor for the protection of the patient and for the diagnostic quality of a radiograph.Moreover, incorrect collimation is one of the main causes for a retake and the associated costs.In this paper we propose a novel collimation optimization approach using Time-of-Flight cameras and deep Neural Networks trained end-to-end to increase the diagnostic quality of a radiograph.For this we acquired a new dataset in a clinical environment consisting of depth images of the lower leg and the abdomen.Using this dataset we are able to segment depth images for the optimal collimation with an average IoU of 83%.* Contributed equally.The order of author names was randomly determined.† We thank Celina Schubbe for her help and support in collecting the dataset.
Dominik Mairhöfer, Manuel Laufer, Lennart Berkel, Arpad Bischof, Erhardt Barth, Jörg Barkhausen, Thomas Martinetz
ESANN5
2024 Video Understanding Using 2D-CNNs on Salient Spatio-Temporal Slices
Yaxin Hu 0001, Erhardt Barth
ICANN (3)2
2024 How to Efficiently Use Color and Temporal Information for Video Understanding
Yaxin Hu 0001, Erhardt Barth
ICONIP (8)2
2024 Novel Design Ideas that Improve Video-Understanding Networks with Transformers
abstract
With the development of deep learning, video understanding has become a promising and challenging research field. In recent years, different transformer architectures have shown state-of-the-art performance on most benchmarks. Although transformers can process longer temporal sequences and therefor perform better than convolution networks, they require huge datasets and have high computational costs. The inputs to video transformers are usually clips sampled out of a video, and the length of the clips is limited by the available computing resources. In this paper, we introduce novel methods to sample and tokenize the input video, such as to better capture the dynamics of the input without a large increase in computational costs. Moreover, we introduce the MinBlocks as a novel architecture inspired by neural processing in biological vision. The combination of variable tubes and MinBlocks improves network performance by 10.67%.
Yaxin Hu 0001, Erhardt Barth
IJCNN2
2024 Leaky ReLUs That Differ in Forward and Backward Pass Facilitate Activation Maximization in Deep Neural Networks
abstract
Activation maximization (AM) strives to generate optimal input stimuli, revealing features that trigger high responses in trained deep neural networks. AM is an important method of explainable AI. We demonstrate that AM fails to produce optimal input stimuli for simple functions containing ReLUs or Leaky ReLUs, casting doubt on the practical usefulness of AM and the visual interpretation of the generated images. This paper proposes a solution based on using Leaky ReLUs with a high negative slope in the backward pass while keeping the original, usually zero, slope in the forward pass. The approach significantly increases the maxima found by AM. The resulting ProxyGrad algorithm implements a novel optimization technique for neural networks that employs a secondary network as a proxy for gradient computation. This proxy network is designed to have a simpler loss landscape with fewer local maxima than the original network. Our chosen proxy network is an identical copy of the original network, including its weights, with distinct negative slopes in the Leaky ReLUs. Moreover, we show that ProxyGrad can be used to train the weights of Convolutional Neural Networks for classification such that, on some of the tested benchmarks, they outperform traditional networks.
Christoph Linse, Erhardt Barth, Thomas Martinetz
IJCNN2
2023 Connections Between Pairs of Filters Improve the Accuracy of Convolutional Neural Networks
abstract
While researchers continue to find new and improved network structures for CNNs, most of the newly invented architectures still rely on the traditional pattern of stacking convolutional blocks and separating them with pointwise activation functions. However, there are drawbacks to a network purely building on pointwise nonlinearities. One alternative is to introduce a pairwise connection between two filters of a network. Typical connection functions use multiplications or the minimum operation to realize logical AND connections. In this paper, we go one step further by demonstrating that CNNs can benefit from more general connections, which include parameters that are learned. With such parameters, the network is able to implement different connections in different network layers and better adapt the connection function to the task at hand.
Kathleen Anderson, Philipp Grüning, Erhardt Barth
IJCNN3
2023 Convolutional Neural Networks Do Work with Pre-Defined Filters
abstract
We present a novel class of Convolutional Neural Networks called Pre-defined Filter Convolutional Neural Networks (PFCNNs), where all$n\times n$convolution kernels with$n > 1$are pre-defined and constant during training. It involves a special form of depthwise convolution operation called a Pre-defined Filter Module (PFM). In the channel-wise convolution part, the$1\times n\times n$kernels are drawn from a fixed pool of only a few (16) different pre-defined kernels. In the$1\times 1$convolution part linear combinations of the pre-defined filter outputs are learned. Despite this harsh restriction, complex and discriminative features are learned. These findings provide a novel perspective on the way how information is processed within deep CNNs. We discuss various properties of PFCNNs and prove their effectiveness using the popular datasets Caltech101, CIFAR10, CUB-200-2011, FGVC-Aircraft, Flowers102, and Stanford Cars. Our implementation of PFCNNs is provided on Github https://github.com/Criscraft/PredefinedFilterNetworks.
Christoph Linse, Erhardt Barth, Thomas Martinetz
IJCNN2
2020 Log-Nets: Logarithmic Feature-Product Layers Yield More Compact Networks
Philipp Grüning, Thomas Martinetz, Erhardt Barth
ICANN (2)3
2020 A Multi-Organ Nucleus Segmentation Challenge
abstract
Generalized nucleus segmentation techniques can contribute greatly to reducing the time to develop and validate visual biomarkers for new digital pathology datasets. We summarize the results of MoNuSeg 2018 Challenge whose objective was to develop generalizable nuclei segmentation techniques in digital pathology. The challenge was an official satellite event of the MICCAI 2018 conference in which 32 teams with more than 80 participants from geographically diverse institutes participated. Contestants were given a training set with 30 images from seven organs with annotations of 21,623 individual nuclei. A test dataset with 14 images taken from seven organs, including two organs that did not appear in the training set was released without annotations. Entries were evaluated based on average aggregated Jaccard index (AJI) on the test set to prioritize accurate instance segmentation as opposed to mere semantic segmentation. More than half the teams that completed the challenge outperformed a previous baseline. Among the trends observed that contributed to increased accuracy were the use of color normalization as well as heavy data augmentation. Additionally, fully convolutional networks inspired by variants of U-Net, FCN, and Mask-RCNN were popularly used, typically based on ResNet or VGG base architectures. Watershed segmentation on predicted semantic segmentation maps was a popular post-processing strategy. Several of the top techniques compared favorably to an individual human annotator and can be used with confidence for nuclear morphometrics.
Neeraj Kumar 0002, Ruchika Verma, Deepak Anand, Yanning Zhou 0001, Omer Fahri Onder, Efstratios Tsougenis, Hao Chen 0011, Pheng-Ann Heng, Jiahui Li 0005, Navid Alemi Koohbanani, Mostafa Jahanifar, Neda Zamani Tajeddin, Ali Gooya, Nasir M. Rajpoot, Xuhua Ren, Sihang Zhou 0001, Qian Wang 0001, Dinggang Shen, Cheng-Kun Yang, Chi-Hung Weng, Wei-Hsiang Yu, Chao-Yuan Yeh, Shuoyu Xu, Pak-Hei Yeung, Amirreza Mahbod, Gerald Schaefer, Isabella Ellinger, Rupert Ecker, Örjan Smedby, Chunliang Wang, Benjamin Chidester, Vinh Ton-That, Minh-Triet Tran, Jian Ma 0004, Minh N. Do, Simon Graham, Quoc Dang Vu, Jin Tae Kwak, Akshaykumar Gunda, Raviteja Chunduri, Corey Hu, Dariush Lotfi, Reza Safdari, Antanas Kascenas, Alison O'Neil, Dennis Eschweiler, Johannes Stegmaier, Yanping Cui, Kailin Chen, Xinmei Tian 0001, Philipp Grüning, Erhardt Barth, Elad Arbel, Itay Remer, Amir Ben-Dor, Ekaterina Sirazitdinova, Matthias Kohl, Stefan Braunewell, Yuexiang Li, Xinpeng Xie, LinLin Shen, Jun Ma 0016, Krishanu Das Baksi, Mohammad Azam Khan, Jaegul Choo, Adrián Colomer, Valery Naranjo, Linmin Pei, Khan M. Iftekharuddin, Kaushiki Roy, Debotosh Bhattacharjee, Aníbal Pedraza, Gloria Bueno García, Sabarinathan Devanathan, Saravanan Radhakrishnan, Praveen Koduganty, Zihan Wu 0001, Guanyu Cai, Amit Sethi
IEEE Trans. Medical Imaging58
2017 Sensing Forest for Pattern Recognition
Irina Burciu, Thomas Martinetz, Erhardt Barth
ACIVS3
2017 A Hybrid Convolutional Variational Autoencoder for Text Generation
abstract
In this paper we explore the effect of architectural choices on learning a variational autoencoder (VAE) for text generation. In contrast to the previously introduced VAE model for text where both the encoder and decoder are RNNs, we propose a novel hybrid architecture that blends fully feed-forward convolutional and deconvolutional components with a recurrent language model. Our architecture exhibits several attractive properties such as faster run time and convergence, ability to better handle long sequences and, more importantly, it helps to avoid the issue of the VAE collapsing to a deterministic model.
Stanislau Semeniuta, Aliaksei Severyn, Erhardt Barth
EMNLP3
2017 Recursive autoconvolution for unsupervised learning of convolutional neural networks
abstract
In visual recognition tasks, such as image classification, unsupervised learning exploits cheap unlabeled data and can help to solve these tasks more efficiently. We show that the recursive autoconvolution operator, adopted from physics, boosts existing unsupervised methods by learning more discriminative filters. We take well established convolutional neural networks and train their filters layer-wise. In addition, based on previous works we design a network which extracts more than 600k features per sample, but with the total number of trainable parameters greatly reduced by introducing shared filters in higher layers. We evaluate our networks on the MNIST, CIFAR-10, CIFAR-100 and STL-10 image classification benchmarks and report several state of the art results among other unsupervised methods.
Boris Knyazev 0001, Erhardt Barth, Thomas Martinetz
IJCNN2
2016 Recurrent Dropout without Memory Loss
abstract
This paper presents a novel approach to recurrent neural network (RNN) regularization. Differently from the widely adopted dropout method, which is applied to forward connections of feedforward architectures or RNNs, we propose to drop neurons directly in recurrent connections in a way that does not cause loss of long-term memory. Our approach is as easy to implement and apply as the regular feed-forward dropout and we demonstrate its effectiveness for the most effective modern recurrent network – Long Short-Term Memory network. Our experiments on three NLP benchmarks show consistent improvements even when combined with conventional feed-forward dropout.
Stanislau Semeniuta, Aliaksei Severyn, Erhardt Barth
COLING3
2015 Deep convolutional neural networks as generic feature extractors
abstract
Recognizing objects in natural images is an intricate problem involving multiple conflicting objectives. Deep convolutional neural networks, trained on large datasets, achieve convincing results and are currently the state-of-the-art approach for this task. However, the long time needed to train such deep networks is a major drawback. We tackled this problem by reusing a previously trained network. For this purpose, we first trained a deep convolutional network on the ILSVRC-12 dataset. We then maintained the learned convolution kernels and only retrained the classification part on different datasets. Using this approach, we achieved an accuracy of 67.68% on CIFAR-100, compared to the previous state-of-the-art result of 65.43%. Furthermore, our findings indicate that convolutional networks are able to learn generic feature extractors that can be used for different tasks.
Lars Hertel, Erhardt Barth, Thomas Käster, Thomas Martinetz
IJCNN2
2015 Learning orthogonal sparse representations by using geodesic flow optimization
abstract
In this paper we propose the novel algorithm GF-OSC, which learns an orthogonal basis that provides an optimal K-sparse data representation for a given set of training samples. The underlying optimization problem is composed of two nested subproblems: (i) given a basis, to determine an optimal K-sparse coefficient vector for each data sample, and (ii) given a K-sparse coefficient vector for each data sample, to determine an optimal basis. Both subproblems have closed form solutions, which can be computed alternately in an iterative manner. Due to the nesting of the subproblems, however, this approach can only find an optimal solution if the underlying sparsity level is sufficiently high. To overcome this shortcoming, our GF-OSC algorithm solves subproblem (ii) via gradient descent on the corresponding cost function within the underlying lower dimensional space of free dictionary parameters. This algorithmic substep is based on the geodesic flow optimization framework proposed by Plumbley. On synthetic data, we show in a comparison with four alternative learning algorithms the superior recovery performance of GF-OSC and show that it needs significantly fewer learning epochs to converge. Furthermore, we demonstrate the potential of GF-OSC for image compression. For five standard test images, we derived sparse image approximations based on a GF-OSC basis that was trained on natural image patches. In terms of PSNR, the approximation performance of the GF-OSC basis is between 0.09 to 0.32 dB higher compared to using the 2D DCT basis, and between 1.66 to 3.4 dB higher compared to using the 2D Haar wavelet basis.
Henry Schütze, Erhardt Barth, Thomas Martinetz
IJCNN2
2015 Self-organizing maps for hand and full body tracking
Foti Coleca, Andreea State, Sascha Klement, Erhardt Barth, Thomas Martinetz
Neurocomputing4
2012 Novelty detection for the inspection of light-emitting diodes
Fabian Timm, Erhardt Barth
Expert Syst. Appl.2
2012 Intrinsic Dimensionality Predicts the Saliency of Natural Dynamic Scenes
abstract
Since visual attention-based computer vision applications have gained popularity, ever more complex, biologically inspired models seem to be needed to predict salient locations (or interest points) in naturalistic scenes. In this paper, we explore how far one can go in predicting eye movements by using only basic signal processing, such as image representations derived from efficient coding principles, and machine learning. To this end, we gradually increase the complexity of a model from simple single-scale saliency maps computed on grayscale videos to spatiotemporal multiscale and multispectral representations. Using a large collection of eye movements on high-resolution videos, supervised learning techniques fine-tune the free parameters whose addition is inevitable with increasing complexity. The proposed model, although very simple, demonstrates significant improvement in predicting salient locations in naturalistic videos over four selected baseline models and two distinct data labeling scenarios.
Eleonora Vig, Michael Dorr, Thomas Martinetz, Erhardt Barth
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Gaze guidance reduces the number of collisions with pedestrians in a driving simulator
abstract
Our study explores the potential of gaze guidance in driving and analyzes eye movements and driving behavior in safety-critical situations. We collected eye movements from subjects instructed to drive predetermined routes in a driving simulator. While driving, the subjects performed various cognitive tasks designed to divert their attention away from the road. The 30 subjects were equally divided in two groups, a control and a gaze guidance group. For the latter, potentially dangerous events, such as a pedestrian suddenly crossing the street, were highlighted with temporally transient gaze-contingent cues, which were triggered if the subject did not look at the pedestrian. For the group that drove with gaze guidance, eye movements have a reduced variability after the gaze-capturing event and shorter reaction times to it. More importantly, gaze guidance leads to a safer driving behavior and a significantly reduced number of collisions.
Laura Pomârjanschi, Michael Dorr, Erhardt Barth
ACM Trans. Interact. Intell. Syst.3
2011 Soft-competitive learning of sparse codes and its application to image reconstruction
Kai Labusch, Erhardt Barth, Thomas Martinetz
Neurocomputing2
2010 Space-variant spatio-temporal filtering of video for gaze visualization and perceptual learning
abstract
We introduce an algorithm for space-variant filtering of video based on a spatio-temporal Laplacian pyramid and use this algorithm to render videos in order to visualize prerecorded eye movements. Spatio-temporal contrast and colour saturation are reduced as a function of distance to the nearest gaze point of regard, i.e. non-fixated, distracting regions are filtered out, whereas fixated image regions remain unchanged. Results of an experiment in which the eye movements of an expert on instructional videos are visualized with this algorithm, so that the gaze of novices is guided to relevant image locations, show that this visualization technique facilitates the novices' perceptual learning.
Michael Dorr, Halszka Jarodzka, Erhardt Barth
ETRA3
2010 A Learned Saliency Predictor for Dynamic Natural Scenes
Eleonora Vig, Michael Dorr, Thomas Martinetz, Erhardt Barth
ICANN (3)4
2010 Time-of-Flight Cameras in Computer Graphics
abstract
Abstract A growing number of applications depend on accurate and fast 3D scene analysis. Examples are model and lightfield acquisition, collision prevention, mixed reality and gesture recognition. The estimation of a range map by image analysis or laser scan techniques is still a time‐consuming and expensive part of such systems. A lower‐priced, fast and robust alternative for distance measurements are time‐of‐flight (ToF) cameras. Recently, significant advances have been made in producing low‐cost and compact ToF devices, which have the potential to revolutionize many fields of research, including computer graphics, computer vision and human machine interaction (HMI). These technologies are starting to have an impact on research and commercial applications. The upcoming generation of ToF sensors, however, will be even more powerful and will have the potential to become ‘ubiquitous real‐time geometry devices’ for gaming, web‐conferencing, and numerous other applications. This paper gives an account of recent developments in ToF technology and discusses the current state of the integration of this technology into various graphics‐related applications.
Andreas Kolb 0001, Erhardt Barth, Reinhard Koch, Rasmus Larsen 0001
Comput. Graph. Forum2
2010 Shading constraint improves accuracy of time-of-flight measurements
Martin Böhme, Martin Haker, Thomas Martinetz, Erhardt Barth
Comput. Vis. Image Underst.4
2010 Special issue on Time-of-Flight camera based computer vision
Rasmus Larsen 0001, Erhardt Barth, Andreas Kolb 0001
Comput. Vis. Image Underst.2
2009 Multimodal Sparse Features for Object Detection
Martin Haker, Thomas Martinetz, Erhardt Barth
ICANN (2)3
2009 Sparse Coding Neural Gas: Learning of overcomplete data representations
Kai Labusch, Erhardt Barth, Thomas Martinetz
Neurocomputing2
2008 Learning Data Representations with Sparse Coding Neural Gas
Kai Labusch, Erhardt Barth, Thomas Martinetz
ESANN2
2008 A software framework for simulating eye trackers
abstract
We describe an open-source software framework that simulates the measurements made using one or several cameras in a video-oculographic eye tracker. The framework can be used to compare objectively the performance of different eye tracking setups (number and placement of cameras and light sources) and gaze estimation algorithms. We demonstrate the utility of the framework by using it to compare two remote eye tracking methods, one using a single camera, the other using two cameras.
Martin Böhme, Michael Dorr, Mathis Graw, Thomas Martinetz, Erhardt Barth
ETRA5
2008 Sparse Coding Neural Gas for the Separation of Noisy Overcomplete Sources
Kai Labusch, Erhardt Barth, Thomas Martinetz
ICANN (1)2
2008 Simple Method for High-Performance Digit Recognition Based on Sparse Coding
abstract
In this brief paper, we propose a method of feature extraction for digit recognition that is inspired by vision research: a sparse-coding strategy and a local maximum operation. We show that our method, despite its simplicity, yields state-of-the-art classification results on a highly competitive digit-recognition benchmark. We first employ the unsupervised Sparsenet algorithm to learn a basis for representing patches of handwritten digit images. We then use this basis to extract local coefficients. In a second step, we apply a local maximum operation to implement local shift invariance. Finally, we train a support vector machine (SVM) on the resulting feature vectors and obtain state-of-the-art classification performance in the digit recognition task defined by the MNIST benchmark. We compare the different classification performances obtained with sparse coding, Gabor wavelets, and principal component analysis (PCA). We conclude that the learning of a sparse representation of local image patches combined with a local maximum operation for feature extraction can significantly improve recognition performance.
Kai Labusch, Erhardt Barth, Thomas Martinetz
IEEE Trans. Neural Networks2
2006 Gaze-contingent temporal filtering of video
abstract
We describe an algorithm for manipulating the temporal resolution of a video in real time, contingent upon the viewer's direction of gaze. The purpose of this work is to study the effect that a controlled manipulation of the temporal frequency content in real-world scenes has on eye movements. We build on the work of Perry and Geisler [1998; 2002], who manipulate spatial resolution as a function of gaze direction, allowing them to mimic the resolution distribution of the human retina or to simulate the effect of various diseases (e.g. glaucoma).Our temporal filtering algorithm is similar to that of Perry and Geisler in that we interpolate between the levels of a multiresolution pyramid. However, in our case, the pyramid is built along the temporal dimension, and this requires careful management of the buffering of video frames and of the order in which the filtering operations are performed. On a standard personal computer, the algorithm achieves real-time performance (30 frames per second) on high-resolution videos (960 by 540 pixels).We present experimental results showing that the manipulation performed by the algorithm reduces the number of high-amplitude saccades and can remain unnoticed by the observer.
Martin Böhme, Michael Dorr, Thomas Martinetz, Erhardt Barth
ETRA4
2006 Eye movement predictions on natural videos
Martin Böhme, Michael Dorr, Christopher Krause, Thomas Martinetz, Erhardt Barth
Neurocomputing5
2006 Analysis of Superimposed Oriented Patterns
abstract
Estimation of local orientation in images may be posed as the problem of finding the minimum gray-level variance axis in a local neighborhood. In bivariate images, the solution is given by the eigenvector corresponding to the smaller eigenvalue of a 2 x 2 tensor. For an ideal single orientation, the tensor is rank-deficient, i.e., the smaller eigenvalue vanishes. A large minimal eigenvalue signals the presence of more than one local orientation, what may be caused by non-opaque additive or opaque occluding objects, crossings, bifurcations, or corners. We describe a framework for estimating such superimposed orientations. Our analysis is based on the eigensystem analysis of suitably extended tensors for both additive and occluding superpositions. Unlike in the single-orientation case, the eigensystem analysis does not directly yield the orientations, rather, it provides so-called mixed-orientation parameters (MOPs). We, therefore, show how to decompose the MOPs into the individual orientations. We also show how to use tensor invariants to increase efficiency, and derive a new feature for describing local neighborhoods which is invariant to rigid transformations. Applications are, e.g., in texture analysis, directional filtering and interpolation, feature extraction for corners and crossings, tracking, and signal separation.
Til Aach, Cicero Mota, Ingo Stuke, Matthias Mühlich, Erhardt Barth
IEEE Trans. Image Process.5
2005 Spatial and spectral analysis of occluded motions
Cicero Mota, Ingo Stuke, Til Aach, Erhardt Barth
Signal Process. Image Commun.4
2004 Estimation of multiple local orientations in image signals
abstract
Local orientation estimation can be posed as the problem of finding the minimum grey level variance axis within a local neighbourhood. In 2D image signals, this corresponds to the eigensystem analysis of a 2 /spl times/ 2-tensor, which yields valid results for single orientations. We describe extensions to multiple overlaid orientations, which may be caused by transparent objects, crossings, bifurcations, corners etc. Multiple orientation detection is based on the eigensystem analysis of an appropriately extended tensor, yielding so-called mixed orientation parameters. These mixed orientation parameters can be regarded as another tensor built from the sought individual orientation parameters. We show how the mixed orientation tensor can be decomposed into the individual orientations by finding the roots of a polynomial. Applications are, e.g., in directional filtering and interpolation, feature extraction for corners or crossings, and signal separation.
Til Aach, Ingo Stuke, Cicero Mota, Erhardt Barth
ICASSP (3)4
2004 Estimation of multiple orientations in multi-dimensional signals
Cicero Mota, Til Aach, Ingo Stuke, Erhardt Barth
ICIP4
2004 Estimation of multiple motions using block matching and Markov random fields
abstract
This paper deals with the problem of estimating multiple motions at points where these motions are overlaid. We present a new approach that is based on block-matching and can deal with both transparent motions and occlusions. We derive a block-matching constraint for an arbitrary number of moving layers. We use this constraint to design a hierarchical algorithm that can distinguish between the occurrence of single, transparent, and occluded motions and can thus select the appropriate local motion model. The algorithm adapts to the amount of noise in the image sequence by use of a statistical confidence test. The algorithm is further extended to deal with very noisy images by using a regularization based on Markov Random Fields. Performance is demonstrated on image sequences synthesized from natural textures with high levels of additive dynamic noise.
Ingo Stuke, Til Aach, Erhardt Barth, Cicero Mota
VCIP3
2003 Linear and regularized solutions for multiple motions
abstract
We extend a novel framework for the estimation of multiple transparent motions to include regularization. We use mixed-motion parameters to obtain linear Euler-Lagrange equations with a regularization term. The equations are solved iteratively for the mixed-motion parameters based on an update rule that is similar to the case of only one motion. The motion parameters are then obtained as the roots of a complex polynomial of a degree that is equal to the number of overlaid motions. An experimental error analysis is performed and reported.
Ingo Stuke, Til Aach, Cicero Mota, Erhardt Barth
ICASSP (3)4
2003 Spatio-temporal motion estimation for transparency and occlusions
abstract
We present a spatio-temporal analysis of motion at occluding boundaries as an extension of previous results for trans- parent motions. We show how these new results generalize alternative approaches derived in the Fourier domain that are limited by assuming straight occlusion boundaries. Furthermore, we derive a novel hierarchical algorithm that can deal with single, multiple-transparent, and occluded motions.
Ingo Stuke, Til Aach, Erhardt Barth, Cicero Mota
ICIP (3)3
2003 Categorization of Transparent-Motion Patterns Using the Projective Plane
Cicero Mota, Michael Dorr, Ingo Stuke, Erhardt Barth
SNPD4
2003 Estimation of Multiple Motions by Block Matching
Ingo Stuke, Til Aach, Erhardt Barth, Cicero Mota
SNPD3
2001 Analytic solutions for multiple motions
abstract
A novel framework for single and multiple motion estimation is presented. It is based on a generalized structure tensor that contains blurred products of directional derivatives. The order of differentiation increases with the number of motions but more general linear filters can be used instead of derivatives. From the general framework, a hierarchical algorithm for motion estimation is derived and its performance is demonstrated on a synthetic sequence.
Cicero Mota, Ingo Stuke, Erhardt Barth
ICIP (2)3
1993 Image Encoding, Labeling, and Reconstruction from Differential Geometry
Erhardt Barth, Terry Caelli, Christoph Zetzsche
CVGIP Graph. Model. Image Process.1
1991 Direct detection of flow discontinuities by 3D curvature operators
Christoph Zetzsche, Erhardt Barth
Pattern Recognit. Lett.2