Lutz Goldmann

dblp:79/2218 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 6 first-authorArtificial intelligence and machine learning · 4 · 1 since 2021Databases, data management, data science and information retrieval · 2Security and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
2 papers
Rendering · 38% Multimedia systems and quality of experience · 29% Image and video processing · 19%
Artificial intelligence
1 paper
Image recognition and object detection · 62% Face, body and person analysis · 38%

Topics — the 10 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Rendering
3d view synthesis
0.212013
Adaptive Image Warping for Hole Prevention in 3D View Synthesis · IEEE Trans. Image Process. 2013
Rendering › image-based rendering
depth-image-based rendering
0.212013
Adaptive Image Warping for Hole Prevention in 3D View Synthesis · IEEE Trans. Image Process. 2013
Image and video processing
image warping
0.212013
Adaptive Image Warping for Hole Prevention in 3D View Synthesis · IEEE Trans. Image Process. 2013
Multimedia systems and quality of experience › subjective quality assessment
pairwise comparison
0.112011
A new analysis method for paired comparison and its application to 3D quality assessment · ACM Multimedia 2011
Image and video coding
quality assessment
0.112011
A new analysis method for paired comparison and its application to 3D quality assessment · ACM Multimedia 2011
Multimedia systems and quality of experience
subjective quality assessment
0.112011
A new analysis method for paired comparison and its application to 3D quality assessment · ACM Multimedia 2011
Computer vision › Image recognition and object detection › object detection
component-based detection
0.112007
Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions · IEEE Trans. Inf. Forensics Secur. 2007
Computer vision › Face, body and person analysis
face detection
0.112007
Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions · IEEE Trans. Inf. Forensics Secur. 2007
Computer vision › Image recognition and object detection
object detection
0.012007
Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions · IEEE Trans. Inf. Forensics Secur. 2007
Computer vision › Image recognition and object detection
structural pattern recognition
0.012007
Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions · IEEE Trans. Inf. Forensics Secur. 2007

Methods — techniques the papers use, named apart from their topics

quadtree decomposition · 0.2optimization · 0.2outlier detection · 0.1maximum likelihood estimation · 0.1haar-like features · 0.1graph matching · 0.1adaboost · 0.1
YearPublicationVenuePosition
2022 Table detection in business document images by message passing networks
Pau Riba, Lutz Goldmann, Oriol Ramos Terrades, Diede Rusticus, Alicia Fornés, Josep Lladós 0001
Pattern Recognit.2
2019 Table Detection in Invoice Documents by Graph Neural Networks
abstract
Tabular structures in documents offer a complementary dimension to the raw textual data, representing logical or quantitative relationships among pieces of information. In digital mail room applications, where a large amount of administrative documents must be processed with reasonable accuracy, the detection and interpretation of tables is crucial. Table recognition has gained interest in document image analysis, in particular in unconstrained formats (absence of rule lines, unknown information of rows and columns). In this work, we propose a graph-based approach for detecting tables in document images. Instead of using the raw content (recognized text), we make use of the location, context and content type, thus it is purely a structure perception approach, not dependent on the language and the quality of the text reading. Our framework makes use of Graph Neural Networks (GNNs) in order to describe the local repetitive structural information of tables in invoice documents. Our proposed model has been experimentally validated in two invoice datasets and achieved encouraging results. Additionally, due to the scarcity of benchmark datasets for this task, we have contributed to the community a novel dataset derived from the RVL-CDIP invoice data. It will be publicly released to facilitate future research.
Pau Riba, Anjan Dutta 0001, Lutz Goldmann, Alicia Fornés, Oriol Ramos Terrades, Josep Lladós 0001
ICDAR3
2019 Document Domain Adaptation with Generative Adversarial Networks
abstract
Despite the digitization of communication flows, information is still commonly exchanged via printed documents. Since manual information extraction becomes inefficient and prone to errors, automatic document processing (ADP) tools have been proposed. Following the current trends in machine learning, many of these tools are based on deep learning methods which require a representative set of training documents to achieve good performance for a target domain. In practice, training documents are often scarce and limited to certain domains which makes it difficult to train a model that generalises well to varying domains. This paper analyses the influence of domain shifts on the performance of document analysis tasks. It further explores the improvements that can be achieved with visual domain adaptation using generative adversarial networks (GANs). The results show that the impact of the domain shift on the performance is not only depending on the difference between the domains but also on the analysis task itself. While some tasks such as document binarization are noticeably affected by the domain shift, other tasks like page classification are less sensitive. It is also shown that the use of mapped training data obtained from a GAN, which translates between the source and target domain, can improve the performance considerably.
Diede Rusticus, Lutz Goldmann, Matthias Reisser, Mauricio Villegas
ICDAR2
2013 Paired comparison-based subjective quality assessment of stereoscopic images
Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
Multim. Tools Appl.2
2013 Adaptive Image Warping for Hole Prevention in 3D View Synthesis
abstract
Increasing popularity of 3D videos calls for new methods to ease the conversion process of existing monocular video to stereoscopic or multi-view video. A popular way to convert video is given by depth image-based rendering methods, in which a depth map that is associated with an image frame is used to generate a virtual view. Because of the lack of knowledge about the 3D structure of a scene and its corresponding texture, the conversion of 2D video, inevitably, however, leads to holes in the resulting 3D image as a result of newly-exposed areas. The conversion process can be altered such that no holes become visible in the resulting 3D view by superimposing a regular grid over the depth map and deforming it. In this paper, an adaptive image warping approach as an improvement to the regular approach is proposed. The new algorithm exploits the smoothness of a typical depth map to reduce the complexity of the underlying optimization problem that is necessary to find the deformation, which is required to prevent holes. This is achieved by splitting a depth map into blocks of homogeneous depth using quadtrees and running the optimization on the resulting adaptive grid. The results show that this approach leads to a considerable reduction of the computational complexity while maintaining the visual quality of the synthesized views.
Nils Plath, Sebastian Knorr, Lutz Goldmann, Thomas Sikora
IEEE Trans. Image Process.3
2012 Geotag propagation in social networks based on user trust model
Peter Vajda, Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
Multim. Tools Appl.4
2011 Social game epitome versus automatic visual analysis
abstract
With the rapid growth of digital photography, sharing of photos with friends and family has become very popular. When people share their photos, they usually organize them in albums according to events or places. To tell the story of some important events in one's life, it is desirable to have an efficient summarization tool which can help people to get a quick overview of an album containing huge number of photos. In this paper, we analyze an approach for photo album summarization through a novel social game “Epitome” as a Facebook application. This social game can collect research data and, at the same time, it provides a collage or a cover photo of the user's photo album, while, at the same time, the user enjoys playing the game. As a benchmark comparison to this game, we performed automatic visual analysis considering several state-of-the-art features.
Peter Vajda, Lutz Goldmann, Touradj Ebrahimi
ICME3
2011 A new analysis method for paired comparison and its application to 3D quality assessment
abstract
Among various subjective quality evaluation methodologies, paired comparison has the advantage of improved simplicity of the subjects' evaluation task due to simplified rating scales and direct comparison of two stimuli. Thus, it may lead to more reliable results when individual quality levels are difficult to define, quality differences between stimuli are small or multiple quality factors are involved. This paper proposes a new method to analyze results of paired comparison-based subjective tests. By assuming that ties convey information about significant differences between two stimuli being compared, the confidence intervals for the quality scores are estimated using a maximum likelihood criterion, which enables us to intuitively examine the significance of quality score differences. We describe the complete test methodology including the test procedure, outlier detection and score analysis applied to quality assessment of 3D images acquired using varying camera distances. Experimental results demonstrate the usefulness of the proposed analysis method, as well as the enhanced quality discriminability of the paired comparison methodology in comparison to the conventional single stimulus methodology.
Jong-Seok Lee, Lutz Goldmann, Touradj Ebrahimi
ACM Multimedia2
2011 Motion parallax based restitution of 3D images on legacy consumer mobile devices
abstract
While 3D display technologies are already widely available for cinema and home or corporate use, only a few portable devices currently feature 3D display capabilities. Moreover, the large majority of 3D display solutions rely on binocular perception. In this paper, we study the alternative methods for restitution of 3D images on conventional 2D displays and analyze their respective performance. This particularly includes the extension of wiggle stereoscopy for portable devices which relies on motion parallax as an additional depth cue. The goal of this paper is to compare two different 3D display techniques, the anaglyph method which provides binocular depth cues and a method based on motion parallax, and to show that the motion parallax based approach to present 3D images on consumer 2D portable screen is an equivalent way in comparison to the above mentioned and well-known anaglyph method. The subsequently conducted subjective quality tests show that viewers even prefer wiggle over anaglyph stereoscopy mainly due to a better color reproduction and a comparable depth perception.
Martin Rerábek, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
MMSP2
2011 Towards high efficiency video coding: Subjective evaluation of potential coding technologies
Francesca De Simone, Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
J. Vis. Commun. Image Represent.2
2010 Temporal synchronization in stereoscopic video: Influence on quality of experience and automatic asynchrony detection
abstract
In this paper, we analyze the influence of temporal asynchrony on the subjective quality of stereoscopic video. Based on our recently created 3D video database, different levels of asynchrony were simulated and a comprehensive subjective test was conducted to determine the associated degradations in quality of experience. Furthermore, we developed a method to detect asynchrony between left and right video streams based on canonical correlation analysis. Experiments demonstrate the robustness of this method with respect to different amounts of asynchrony and scene depth, which makes it suitable to predict quality of experience or automatic resynchronization.
Lutz Goldmann, Jong-Seok Lee, Touradj Ebrahimi
ICIP1
2009 Analysis of the Limits of Graph-Based Object Duplicate Detection
abstract
Several applications require accurate and efficient object duplicate detection methods, such as automatic video and image tag propagation, video surveillance, and high level image or video search. In this paper, we explore the limits of our recently proposed graph-based object duplicate detection method. The dependency of the performance with respect to the number of training images is assessed and the optimal detection parameters are determined. Furthermore, the differences among various object classes are analyzed. In this way, this paper provides an in-depth analysis of the graph based object duplicate detection method.
Peter Vajda, Lutz Goldmann, Touradj Ebrahimi
ISM2
2009 Multimodal person search combining information fusion and relevance feedback
abstract
With the increasing amount of multimedia data, efficient tools for search and retrieval are needed. Since people are naturally one of the most interesting objects within these documents, a system for multimodal person search and retrieval has been developed. It combines the audiovisual analysis of persons with the query by example paradigm and relevance feedback to provide an efficient tool for searching multimedia data. For the relevance feedback, one and two class approaches are considered and compared to each other. Multimodal fusion techniques are used to exploit the complementary character of the audio and video information. The experimental results prove that multimodal person search and retrieval is feasible and more efficient than manual exploration.
Lutz Goldmann, Amjad Samour, Touradj Ebrahimi, Thomas Sikora
MMSP1
2008 Towards Fully Automatic Image Segmentation Evaluation
Lutz Goldmann, Tomasz Adamek, Peter Vajda, Mustafa Karaman, Roland Mörzinger, Eric Galmar, Thomas Sikora, Noel E. O'Connor, Thien Ha-Minh, Touradj Ebrahimi, Peter Schallauer, Benoit Huet
ACIVS1
2008 More robust face recognition by considering occlusion information
abstract
This paper addresses one of the main challenges of face recognition (FR): facial occlusions. Currently, the human brain is the most robust known FR approach towards partially occluded faces. Nevertheless, it is still not clear if humans recognize faces using a holistic or a component-based strategy, or even a combination of both. In this paper, three different approaches based on principal component analysis (PCA) are analyzed. The first one, a holistic approach, is the well-known eigenface approach. The second one, a component-based method, is a variation of the eigenfeatures approach, and finally, the third one, a near-holistic method, is an extension of the lophoscopic principal component analysis (LPCA). So the main contributions of this paper are: The three different strategies are compared and analyzed for identifying partially occluded faces and furthermore it explores how a priori knowledge about present occlusions can be used to improve the recognition performance.
Antonio Rama, Francesc Tarres, Lutz Goldmann, Thomas Sikora
FG3
2008 On the detection and localization of facial occlusions and its use within different scenarios
abstract
Face analysis is a very active research field, due to its large variety of applications and the different challenges (illumination, pose, expressions or occlusions) the methods need to cope with. Facial occlusions are one of the biggest challenges since they are difficult to model and have a large influence on the performance of subsequent analysis modules. This paper describes a face detection/classification module that allows to detect and localize faces and present occlusions and discusses the use of this additional information within different application scenarios. The approach is evaluated on two databases with realistic occlusions and performs very well for the different detection/classification tasks. It achieves a f-measure of over 97% for face detection and around 86% for component detection. Regarding the occlusion detection, the proposed approach reaches a recognition rate above 91% for both faces and components.
Lutz Goldmann, Antonio Rama, Thomas Sikora, Francesc Tarres
MMSP1
2007 Components and Their Topology for Robust Face Detection in the Presence of Partial Occlusions
abstract
This paper presents a novel approach for automatic and robust object detection. It utilizes a component-based approach that combines techniques from both statistical and structural pattern recognition domain. While the component detection relies on Haar-like features and an AdaBoost trained classifier cascade, the topology verification is based on graph matching techniques. The system was applied to face detection and the experiments show its outstanding performance in comparison to conventional face detection approaches. Especially in the presence of partial occlusions, uneven illumination, and out-of-plane rotations, it yields higher robustness. Furthermore, this paper provides a comprehensive review of recent approaches for object detection and gives an overview of available databases for face detection.
Lutz Goldmann, Ullrich J. Mönich, Thomas Sikora
IEEE Trans. Inf. Forensics Secur.1
2006 Extracting High Level Semantics by Means of Speech, Audio, and Image Primitives in Surveillance Applications
abstract
Traditional surveillance systems are usually based on visual information only. With the emerging multimedia analysis techniques, interests are changing towards systems that incorporate multiple sensors and different modalities, which leads to new ways of analyzing this multimedia data and more sophisticated applications. This paper shortly reviews the ideas of traditional surveillance systems and explains actual research interests in this domain. Then, it focuses on the typical structure, goals, and applications of multimedia surveillance systems. These issues are supported by short descriptions of selected analysis steps of such a system currently under development. Some experimental results are given to illustrate the extracted semantics and to assess the performance of the individual steps.
Lutz Goldmann, Amjad Samour, Mustafa Karaman, Thomas Sikora
ICIP1
2004 Human body posture recognition using MPEG-7 descriptors
abstract
This paper presents a novel approach to human body posture recognition based on the MPEG-7 contour-based shape descriptor and the widely used projection histogram. A combination of them was used to recognize the main posture and the view of a human based on the binary object mask obtained by the segmentation process. The recognition is treated as a typical pattern recognition task and is carried out through a hierarchy of classifiers. Therefore various structures both hierachical and non-hierarchical, in combination with different classifiers, are compared to each other with respect to recognition performance and computational complexity. Based on this an optimal system design with recognition rates of 95.59% for the main posture, 77.84% for the view and 79.77% in combination is achieved.
Lutz Goldmann, Mustafa Karaman, Thomas Sikora
VCIP1