Matthias Zeppelzauer

dblp:81/4405 · DBLP profile ↗
← Back
36ranked-venue papers
8as first author
15since 2021 · last 2026
0000-0003-0413-4746ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 24 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 11 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Computer networks · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Distilling knowledge from large language models: A concept bottleneck model for hate and counter speech recognition
abstract
The rapid increase in hate speech on social media has exposed an unprecedented impact on society, making automated methods for detecting such content important. Unlike prior black-box models, we propose a novel transparent method for automated hate and counter speech recognition, i.e., “Speech Concept Bottleneck Model” (SCBM), using adjectives as human-interpretable bottleneck concepts. SCBM leverages large language models (LLMs) to map input texts to an abstract adjective-based representation, which is then sent to a light-weight classifier for downstream tasks. Across five benchmark datasets spanning multiple languages and platforms (e.g., Twitter, Reddit, YouTube), SCBM achieves an average macro-F1 score of 0.69 which outperforms the most recently reported results from the literature on four out of five datasets. Aside from high recognition accuracy, SCBM provides a high level of both local and global interpretability. Furthermore, fusing our adjective-based concept representation with transformer embeddings, leads to a 1.8% performance increase on average across all datasets, showing that the proposed representation captures complementary information. Our results demonstrate that adjective-based concept representations can serve as compact, interpretable, and effective encodings for hate and counter speech recognition. With adapted adjectives, our method can also be applied to other NLP tasks.
Roberto Labadie, Djordje Slijepcevic, Xihui Chen, Adrian Jaques Böck, Andreas Babic, Liz Freimann, Christiane Atzmüller, Matthias Zeppelzauer
Inf. Process. Manag.8
2025 Explanatory Interactive Machine Learning for Bias Mitigation in Visual Gender Classification
abstract
Explanatory interactive learning (XIL) enables users to guide model training in machine learning (ML) by providing feedback on the model's explanations, thereby helping it to focus on features that are relevant to the prediction from the user's perspective. In this study, we explore the capability of this learning paradigm to mitigate bias and spurious correlations in visual classifiers, specifically in scenarios prone to data bias, such as gender classification. We investigate two methodologically different state-of-the-art XIL strategies, i.e., CAIPI and Right for the Right Reasons (RRR), as well as a novel hybrid approach that combines both strategies. The results are evaluated quan-titatively by comparing segmentation masks with explanations generated using Gradient-weighted Class Activation Mapping (GradCAM) and Bounded Logit Attention (BLA). Experimental results demonstrate the effectiveness of these methods in (i) guiding ML models to focus on relevant image features, par-ticularly when CAIPI is used, and (ii) reducing model bias (i.e., balancing the misclassification rates between male and female predictions). Our analysis further supports the potential of XIL methods to improve fairness in gender classifiers. Overall, the increased transparency and fairness obtained by XIL leads to slight performance decreases with an exception being CAIPI, which shows potential to even improve classification accuracy.
Nathanya Queby Satriani, Djordje Slijepcevic, Markus Schedl, Matthias Zeppelzauer
CBMI4
2025 CounterHelp: Promoting Online Civil Courage Among Young People Through AI-Generated Counterspeech
abstract
We present CounterHelp, a mobile app designed to support adolescents and young people in generating customized counter speech in response to hateful comments on TikTok. CounterHelp allows users to specify counter strategies to combat hate speech. After users share a hateful comment with CounterHelp, it retrieves TikTok metadata to capture its context. Leveraging large language models, CounterHelp generates customised and context-sensitive counter speech in one of four predefined styles, particularly tailored for adolescents. A user experience lab study confirms the effectiveness and usability of CounterHelp and the generated counter speech.
Andreas Babic, Xihui Chen, Djordje Slijepcevic, Adrian Jaques Böck, Matthias Zeppelzauer
ACM Multimedia5
2025 A Matter of Time: Revealing the Structure of Time in Vision-Language Models
abstract
Large-scale vision-language models (VLMs) such as CLIP have gained popularity for their generalizable and expressive multimodal representations. By leveraging large-scale training data with diverse textual metadata, VLMs acquire open-vocabulary capabilities, solving tasks beyond their training scope. This paper investigates the temporal awareness of VLMs, assessing their ability to position visual content in time. We introduce TIME10k, a benchmark dataset of over 10,000 images with temporal ground truth, and evaluate the time-awareness of 37 VLMs by a novel methodology. Our investigation reveals that temporal information is structured along a low-dimensional, non-linear manifold in the VLM embedding space. Based on this insight, we propose methods to derive an explicit ``timeline'' representation from the embedding space. These representations model time and its chronological progression and thereby facilitate temporal reasoning tasks. Our timeline approaches achieve competitive to superior accuracy compared to a prompt-based baseline while being computationally efficient. All code and data are available at https://tekayanidham.github.io/timeline-page/.
Nidham Tekaya, Manuela Waldner, Matthias Zeppelzauer
ACM Multimedia3
2025 Scalable Class-Centric Visual Interactive Labeling
abstract
Large unlabeled datasets demand efficient and scalable data labeling solutions, in particular when the number of instances and classes is large. This leads to significant visual scalability challenges and imposes a high cognitive load on the users. Traditional instance-centric labeling methods, where (single) instances are labeled in each iteration struggle to scale effectively in these scenarios. To address these challenges, we introduce cVIL, a Class-Centric Visual Interactive Labeling methodology designed for interactive visual data labeling. By shifting the paradigm from assigning-classes-to-instances to assigning-instances-to-classes , cVIL reduces labeling effort and enhances efficiency for annotators working with large, complex and class-rich datasets. We propose a novel visual analytics labeling interface built on top of the conceptual cVIL workflow, enabling improved scalability over traditional visual labeling. In a user study, we demonstrate that cVIL can improve labeling efficiency and user satisfaction over instance-centric interfaces. The effectiveness of cVIL is further demonstrated through a usage scenario, showcasing its potential to alleviate cognitive load and support experts in managing extensive labeling tasks efficiently.
Matthias Matt, Jana Sedlakova, Jürgen Bernard, Matthias Zeppelzauer, Manuela Waldner
Comput. Graph.4
2025 Interactive Discovery and Exploration of Visual Bias in Generative Text-to-Image Models
abstract
Abstract Bias in generative Text‐to‐Image (T2I) models is a known issue, yet systematically analyzing such models' outputs to uncover it remains challenging. We introduce the Visual Bias Explorer (ViBEx) to interactively explore the output space of T2I models to support the discovery of visual bias. ViBEx introduces a novel flexible prompting tree interface in combination with zero‐shot bias probing using CLIP for quick and approximate bias exploration. It additionally supports in‐depth confirmatory bias analysis through visual inspection of forward, intersectional, and inverse bias queries. ViBEx is model‐agnostic and publicly available. In four case study interviews, experts in AI and ethics were able to discover visual biases that have so far not been described in literature.
Johannes Eschner, Roberto Labadie, Matthias Zeppelzauer, Manuela Waldner
Comput. Graph. Forum3
2024 Exploring the Plausibility of Hate and Counter Speech Detectors With Explainable Ai
abstract
In this paper we investigate the explainability of transformer models and their plausibility for hate speech and counter speech detection. We compare representatives of four different explainability approaches, i.e., gradient-based, perturbation-based, attention-based, and prototype-based approaches, and analyze them quantitatively with an ablation study and qualitatively in a user study. Results show that perturbationbased explainability performs best, followed by gradient-based and attention-based explainability. Prototype-based experiments did not yield useful results. Overall, we observe that explainability strongly supports the users in better understanding the model predictions.
Adrian Jaques Böck, Djordje Slijepcevic, Matthias Zeppelzauer
CBMI3
2022 Explaining Machine Learning Models for Clinical Gait Analysis
abstract
Machine Learning (ML) is increasingly used to support decision-making in the healthcare sector. While ML approaches provide promising results with regard to their classification performance, most share a central limitation, their black-box character. This article investigates the usefulness of Explainable Artificial Intelligence (XAI) methods to increase transparency in automated clinical gait classification based on time series. For this purpose, predictions of state-of-the-art classification methods are explained with a XAI method called Layer-wise Relevance Propagation (LRP). Our main contribution is an approach that explains class-specific characteristics learned by ML models that are trained for gait classification. We investigate several gait classification tasks and employ different classification methods, i.e., Convolutional Neural Network, Support Vector Machine, and Multi-layer Perceptron. We propose to evaluate the obtained explanations with two complementary approaches: a statistical analysis of the underlying data using Statistical Parametric Mapping and a qualitative evaluation by two clinical experts. A gait dataset comprising ground reaction force measurements from 132 patients with different lower-body gait disorders and 62 healthy controls is utilized. Our experiments show that explanations obtained by LRP exhibit promising statistical properties concerning inter-class discriminativity and are also in line with clinically relevant biomechanical gait characteristics.
Djordje Slijepcevic, Fabian Horst, Sebastian Lapuschkin, Brian Horsak, Anna-Maria Raberger, Andreas Kranzl, Wojciech Samek, Christian Breiteneder, Wolfgang Immanuel Schöllhorn, Matthias Zeppelzauer
ACM Trans. Comput. Heal.10
2022 Machine unlearning: linear filtration for logit-based classifiers
abstract
Abstract Recently enacted legislation grants individuals certain rights to decide in what fashion their personal data may be used and in particular a “right to be forgotten”. This poses a challenge to machine learning: how to proceed when an individual retracts permission to use data which has been part of the training process of a model? From this question emerges the field of machine unlearning , which could be broadly described as the investigation of how to “delete training data from models”. Our work complements this direction of research for the specific setting of class-wide deletion requests for classification models (e.g. deep neural networks). As a first step, we propose linear filtration as an intuitive, computationally efficient sanitization method. Our experiments demonstrate benefits in an adversarial setting over naive deletion schemes.
Thomas Baumhauer, Pascal Schöttle, Matthias Zeppelzauer
Mach. Learn.3
2021 Multimodal Detection of Information Disorder from Social Media
abstract
Social media is accompanied by an increasing pro-portion of content that provides fake information or misleading content, known as information disorder. In this paper, we study the problem of multimodal fake news detection on a large-scale multimodal dataset. We propose a multimodal network architecture that enables different levels and types of information fusion. In addition to the textual and visual content of a posting, we further leverage secondary information, i.e. user comments and metadata. We fuse information at multiple levels to account for the specific intrinsic structure of the modalities. Our results show that multimodal analysis is highly effective for the task and all modalities contribute positively when fused properly.
Matthias Zeppelzauer, Djordje Slijepcevic, Armin Kirchknopf
CBMI1
2021 Where do university graduates live? - A computer vision approach using satellite images
abstract
Abstract In this article, we examine to what extent the settlement of university graduates can be derived from satellite images. We apply a convolutional neural network (CNN) to grid images of a city and predict five density classes of university graduates at a micro level (250 m × 250 m grid size). The CNN reaches an accuracy rate of 40.5% (random approach: 20%). Furthermore, the accuracy increases to 78.3% when considering a one-class deviation compared to the true class. We also examine the predictability of inhabited and uninhabited grid cells, where we achieve an accuracy of 95.3% using the same CNN. From this, we conclude that there is information that correlates with graduate density that can be derived by analysing only satellite images. The findings show the high potential of computer vision for urban and regional economics. Particularly in data-poor regions, the approach utilised facilitates comparative analytics and provides a possible solution for the modifiable aerial unit (MAU) problem. The MAU problem is a statistical bias that can influence the results of a spatial data analysis of point-estimate data that is aggregated in districts of different shapes and sizes, distorting the results.
David Koch, Miroslav Despotovic, Simon Thaler, Matthias Zeppelzauer
Appl. Intell.4
2021 Correction to: Where do university graduates live? - a computer vision approach using satellite images
David Koch, Miroslav Despotovic, Simon Thaler, Matthias Zeppelzauer
Appl. Intell.4
2021 ProSeCo: Visual analysis of class separation measures and dataset characteristics
abstract
Class separation is an important concept in machine learning and visual analytics. We address the visual analysis of class separation measures for both high-dimensional data and its corresponding projections into 2D through dimensionality reduction (DR) methods. Although a plethora of separation measures have been proposed, it is difficult to compare class separation between multiple datasets with different characteristics, multiple separation measures, and multiple DR methods. We present ProSeCo, an interactive visualization approach to support comparison between up to 20 class separation measures and up to 4 DR methods, with respect to any of 7 dataset characteristics: dataset size, dataset dimensions, class counts, class size variability, class size skewness, outlieriness, and real-world vs. synthetically generated data. ProSeCo supports (1) comparing across measures, (2) comparing high-dimensional to dimensionally-reduced 2D data across measures, (3) comparing between different DR methods across measures, (4) partitioning with respect to a dataset characteristic, (5) comparing partitions for a selected characteristic across measures, and (6) inspecting individual datasets in detail. We demonstrate the utility of ProSeCo in two usage scenarios, using datasets [1] posted at https://osf.io/epcf9/.
Jürgen Bernard, Marco Hutter 0002, Matthias Zeppelzauer, Michael Sedlmair, Tamara Munzner
Comput. Graph.3
2021 k-Anonymity in practice: How generalisation and suppression affect machine learning classifiers
abstract
The protection of private information is a crucial issue in data-driven research and business contexts. Typically, techniques like anonymisation or (selective) deletion are introduced in order to allow data sharing, e. g. in the case of collaborative research endeavours. For use with anonymisation techniques, the k-anonymity criterion is one of the most popular, with numerous scientific publications on different algorithms and metrics. Anonymisation techniques often require changing the data and thus necessarily affect the results of machine learning models trained on the underlying data. In this work, we conduct a systematic comparison and detailed investigation into the effects of different k-anonymisation algorithms on the results of machine learning models. We investigate a set of popular k-anonymisation algorithms with different classifiers and evaluate them on different real-world datasets. Our systematic evaluation shows that with an increasingly strong k-anonymity constraint, the classification performance generally degrades, but to varying degrees and strongly depending on the dataset and anonymisation method. Furthermore, Mondrian can be considered as the method with the most appealing properties for subsequent classification.
Djordje Slijepcevic, Maximilian Henzl, Lukas Daniel Klausner, Tobias Dam, Peter Kieseberg, Matthias Zeppelzauer
Comput. Secur.6
2021 A Taxonomy of Property Measures to Unify Active Learning and Human-centered Approaches to Data Labeling
abstract
Strategies for selecting the next data instance to label, in service of generating labeled data for machine learning, have been considered separately in the machine learning literature on active learning and in the visual analytics literature on human-centered approaches. We propose a unified design space for instance selection strategies to support detailed and fine-grained analysis covering both of these perspectives. We identify a concise set of 15 properties, namely measureable characteristics of datasets or of machine learning models applied to them, that cover most of the strategies in these literatures. To quantify these properties, we introduce Property Measures (PM) as fine-grained building blocks that can be used to formalize instance selection strategies. In addition, we present a taxonomy of PMs to support the description, evaluation, and generation of PMs across four dimensions: machine learning (ML) Model Output , Instance Relations , Measure Functionality , and Measure Valence . We also create computational infrastructure to support qualitative visual data analysis: a visual analytics explainer for PMs built around an implementation of PMs using cascades of eight atomic functions. It supports eight analysis tasks, covering the analysis of datasets and ML models using visual comparison within and between PMs and groups of PMs, and over time during the interactive labeling process. We iteratively refined the PM taxonomy, the explainer, and the task abstraction in parallel with each other during a two-year formative process, and show evidence of their utility through a summative evaluation with the same infrastructure. This research builds a formal baseline for the better understanding of the commonalities and differences of instance selection strategies, which can serve as the stepping stone for the synthesis of novel strategies in future work.
Jürgen Bernard, Marco Hutter 0002, Michael Sedlmair, Matthias Zeppelzauer, Tamara Munzner
ACM Trans. Interact. Intell. Syst.4
2019 Towards Distinction of Rock Art Pecking Styles with a Hybrid 2D/3D Approach
abstract
We present a qualitative study on the distinction of pecking styles in prehistoric rock art captured by high-resolution 3D scans. Different pecking styles result from different shapes, sizes, depths and spatial distributions of individual peck marks. Pecking style distinction enables inter alia automated detection of superimpositions. To our knowledge, this is the first attempt towards an automatic analysis and characterization of pecking styles in rock art. We model pecking style similarity by local descriptions of the surface joining full 3D and 2D (image-space) representations. Our results show that different pecking styles can be efficiently retrieved and distinguished by combined 3D/2D surface analysis.
Markus Seidl, Matthias Zeppelzauer
CBMI2
2019 Persistence Bag-of-Words for Topological Data Analysis
abstract
Persistent homology (PH) is a rigorous mathematical theory that provides a robust descriptor of data in the form of persistence diagrams (PDs). PDs exhibit, however, complex structure and are difficult to integrate in today's machine learning workflows. This paper introduces persistence bag-of-words: a novel and stable vectorized representation of PDs that enables the seamless integration with machine learning. Comprehensive experiments show that the new representation achieves state-of-the-art performance and beyond in much less time than alternative approaches.
Bartosz Zielinski 0001, Michal Lipinski, Mateusz Juda, Matthias Zeppelzauer, Pawel Dlotko
IJCAI4
2019 KAVAGait: Knowledge-Assisted Visual Analytics for Clinical Gait Analysis
abstract
In 2014, more than 10 million people in the US were affected by an ambulatory disability. Thus, gait rehabilitation is a crucial part of health care systems. The quantification of human locomotion enables clinicians to describe and analyze a patient's gait performance in detail and allows them to base clinical decisions on objective data. These assessments generate a vast amount of complex data which need to be interpreted in a short time period. We conducted a design study in cooperation with gait analysis experts to develop a novel Knowledge-Assisted Visual Analytics solution for clinical Gait analysis (KAVAGait). KAVAGait allows the clinician to store and inspect complex data derived during clinical gait analysis. The system incorporates innovative and interactive visual interface concepts, which were developed based on the needs of clinicians. Additionally, an explicit knowledge store (EKS) allows externalization and storage of implicit knowledge from clinicians. It makes this information available for others, supporting the process of data inspection and clinical decision making. We validated our system by conducting expert reviews, a user study, and a case study. Results suggest that KAVAGait is able to support a clinician during clinical practice by visualizing complex gait data and providing knowledge of other clinicians.
Markus Wagner 0008, Djordje Slijepcevic, Brian Horsak, Alexander Rind, Matthias Zeppelzauer, Wolfgang Aigner
IEEE Trans. Vis. Comput. Graph.5
2018 Automatic Prediction of Building Age from Photographs
abstract
We present a first method for the automated age estimation of buildings from unconstrained photographs. To this end, we propose a two-stage approach that firstly learns characteristic visual patterns for different building epochs at patch-level and then globally aggregates patch-level age estimates over the building. We compile evaluation datasets from different sources and perform an detailed evaluation of our approach, its sensitivity to parameters, and the capabilities of the employed deep networks to learn characteristic visual age-related patterns. Results show that our approach is able to estimate building age at a surprisingly high level that even outperforms human evaluators and thereby sets a new performance baseline. This work represents a first step towards the automated assessment of building parameters for automated price prediction.
Matthias Zeppelzauer, Miroslav Despotovic, Muntaha Sakeena, David Koch, Mario Döller
ICMR1
2018 SoniControl - A Mobile Ultrasonic Firewall
abstract
The exchange of data between mobile devices in the near-ultrasonic frequency band is a new promising technology for near field communication (NFC) but also raises a number of privacy concerns. We present the first ultrasonic firewall that reliably detects ultrasonic communication and provides the user with effective means to prevent hidden data exchange. This demonstration showcases a new media-based communication technology ("data over audio") together with its related privacy concerns. It enables users to (i) interactively test out and experience ultrasonic information exchange and (ii) shows how to protect oneself against unwanted tracking.
Matthias Zeppelzauer, Alexis Ringot, Florian Taurer
ACM Multimedia1
2018 Towards User-Centered Active Learning Algorithms
abstract
Abstract The labeling of data sets is a time‐consuming task, which is, however, an important prerequisite for machine learning and visual analytics. Visual‐interactive labeling (VIAL) provides users an active role in the process of labeling, with the goal to combine the potentials of humans and machines to make labeling more efficient. Recent experiments showed that users apply different strategies when selecting instances for labeling with visual‐interactive interfaces. In this paper, we contribute a systematic quantitative analysis of such user strategies. We identify computational building blocks of user strategies, formalize them, and investigate their potentials for different machine learning tasks in systematic experiments. The core insights of our experiments are as follows. First, we identified that particular user strategies can be used to considerably mitigate the bootstrap (cold start) problem in early labeling phases. Second, we observed that they have the potential to outperform existing active learning strategies in later phases. Third, we analyzed the identified core building blocks, which can serve as the basis for novel selection strategies. Overall, we observed that data‐based user strategies (clusters, dense areas) work considerably well in early phases, while model‐based user strategies (e.g., class separation) perform better during later phases. The insights gained from this work can be applied to develop novel active learning approaches as well as to better guide users in visual interactive labeling.
Jürgen Bernard, Matthias Zeppelzauer, Markus Lehmann, Michael Sedlmair
Comput. Graph. Forum2
2018 A study on topological descriptors for the analysis of 3D surface texture
Matthias Zeppelzauer, Bartosz Zielinski 0001, Mateusz Juda, Markus Seidl
Comput. Vis. Image Underst.1
2018 Automatic Classification of Functional Gait Disorders
abstract
This paper proposes a comprehensive investigation of the automatic classification of functional gait disorders (GDs) based solely on ground reaction force (GRF) measurements. The aim of this study is twofold: first, to investigate the suitability of the state-of-the-art GRF parameterization techniques (representations) for the discrimination of functional GDs; and second, to provide a first performance baseline for the automated classification of functional GDs for a large-scale dataset. The utilized database comprises GRF measurements from 279 patients with GDs and data from 161 healthy controls (N). Patients were manually classified into four classes with different functional impairments associated with the "hip", "knee", "ankle", and "calcaneus". Different parameterizations are investigated: GRF parameters, global principal component analysis (PCA) based representations, and a combined representation applying PCA on GRF parameters. The discriminative power of each parameterization for different classes is investigated by linear discriminant analysis. Based on this analysis, two classification experiments are pursued: distinction between healthy and impaired gait (N versus GD) and multiclass classification between healthy gait and all four GD classes. Experiments show promising results and reveal among others that several factors, such as imbalanced class cardinalities and varying numbers of measurement sessions per patient, have a strong impact on the classification accuracy and therefore need to be taken into account. The results represent a promising first step toward the automated classification of GDs and a first performance baseline for future developments in this direction.
Djordje Slijepcevic, Matthias Zeppelzauer, Anna-Maria Gorgas, Caterine Schwab, Michael Schüller, Arnold Baca, Christian Breiteneder, Brian Horsak
IEEE J. Biomed. Health Informatics2
2018 Comparing Visual-Interactive Labeling with Active Learning: An Experimental Study
abstract
Labeling data instances is an important task in machine learning and visual analytics. Both fields provide a broad set of labeling strategies, whereby machine learning (and in particular active learning) follows a rather model-centered approach and visual analytics employs rather user-centered approaches (visual-interactive labeling). Both approaches have individual strengths and weaknesses. In this work, we conduct an experiment with three parts to assess and compare the performance of these different labeling strategies. In our study, we (1) identify different visual labeling strategies for user-centered labeling, (2) investigate strengths and weaknesses of labeling strategies for different labeling tasks and task complexities, and (3) shed light on the effect of using different visual encodings to guide the visual-interactive labeling process. We further compare labeling of single versus multiple instances at a time, and quantify the impact on efficiency. We systematically compare the performance of visual interactive labeling with that of active learning. Our main findings are that visual-interactive labeling can outperform active learning, given the condition that dimension reduction separates well the class distributions. Moreover, using dimension reduction in combination with additional visual encodings that expose the internal state of the learning model turns out to improve the performance of visual-interactive labeling.
Jürgen Bernard, Marco Hutter 0002, Matthias Zeppelzauer, Dieter W. Fellner, Michael Sedlmair
IEEE Trans. Vis. Comput. Graph.3
2018 VIAL: a unified process for visual interactive labeling
Jürgen Bernard, Matthias Zeppelzauer, Michael Sedlmair, Wolfgang Aigner
Vis. Comput.2
2017 A study on skeletonization of complex petroglyph shapes
abstract
In this paper, we present a study on skeletonization of real-world shape data. The data stem from the cultural heritage domain and represent contact tracings of prehistoric petroglyphs. Automated analysis can support the work of archeologists on the investigation and categorization of petroglyphs. One strategy to describe petroglyph shapes is skeleton-based. The skeletonization of petroglyphs is challenging since their shapes are complex, contain numerous holes and are often incomplete or disconnected. Thus they pose an interesting testbed for skeletonization. We present a large real-world dataset consisting of more than 1100 petroglyph shapes. We investigate their properties and requirements for the purpose of skeletonization, and evaluate the applicability of state-of-the-art skeletonization and skeleton pruning algorithms on this type of data. Experiments show that pre-processing of the shapes is crucial to obtain robust skeletons. We propose an adaptive pre-processing method for petroglyph shapes and improve several state-of-the-art skeletonization algorithms to make them suitable for the complex material. Evaluations on our dataset show that 79.8 % of all shapes can be improved by the proposed pre-processing techniques and are thus better suited for subsequent skeletonization. Furthermore we observe that a thinning of the shapes produces robust skeletons for 83.5 % of our shapes and outperforms more sophisticated skeletonization techniques.
Ewald Wieser, Markus Seidl, Matthias Zeppelzauer
Multim. Tools Appl.3
2016 Multimodal classification of events in social media
Matthias Zeppelzauer, Daniel Schopfhauser
Image Vis. Comput.1
2015 Efficient image-space extraction and representation of 3D surface topography
abstract
Surface topography refers to the geometric micro-structure of a surface and defines its tactile characteristics (typically in the sub-millimeter range). High-resolution 3D scanning techniques developed recently enable the 3D reconstruction of surfaces including their surface topography. In this paper, we present an efficient image-space technique for the extraction of surface topography from high-resolution 3D reconstructions. Additionally, we filter noise and enhance topographic attributes to obtain an improved representation for subsequent topography classification. Comprehensive experiments show that our representation captures topographic attributes well and significantly improves classification performance compared to alternative 2D and 3D representations.
Matthias Zeppelzauer, Markus Seidl
ICIP1
2015 Social Event Mining in Large Photo Collections
abstract
A significant part of publicly available photos on the Internet depicts a variety of different social events. In order to organize this steadily growing media content and to make it easily accessible, novel indexing methods are required. Essential research questions in this context concern the efficient detection (clustering), classification, and retrieval of social events in large media collections. In this paper we explore two aspects of social events mining. First, the initial clustering of a given photo collection into single events and, second, the retrieval of relevant social events based on user queries. For both aspects we employ commonly available metadata information, such as user, time, GPS data, and user-generated textual descriptions. Performed evaluations in the context of social event detection demonstrate the strong generalization ability of our approach and the potential of contextual data such as time, user, and location. Experiments with social event retrieval clearly indicate the open challenge of mapping between previously detected event clusters and heterogeneous user queries.
Maia Rohm, Matthias Zeppelzauer, Manfred del Fabro, Daniel Schopfhauser
ICMR2
2014 Unsupervised Selection of Robust Audio Feature Subsets
abstract
Feature selection is applied to identify relevant and complementary features from a given high-dimensional feature set. In general, existing filter-based approaches operate on single (scalar) feature components and ignore the relationships among components of multidimensional features. As a result, generated feature subsets lack in interpretability and hardly provide insights into the underlying data. We propose an unsupervised, filter-based feature selection approach that preserves the natural assignment of feature components to semantically meaningful features. Experiments on different tasks in the audio domain show that the proposed approach outperforms well-established feature selection methods in terms of retrieval performance and runtime. Results achieved on different audio datasets for the same retrieval task indicate that the proposed method is more robust in selecting consistent feature sets across different datasets than compared approaches.
Gerhard Sageder, Maia Rohm, Matthias Zeppelzauer
SDM3
2013 Automated social event detection in large photo collections
abstract
The detection of a specific social event requires for high semantic understanding in the interpretation of particular event characteristics such as its type and location. In many cases, photos capturing different events at the same (or highly similar) locations can hardly be distinguished by each other. Available metadata can provide assistance where there is no expert knowledge at hand. However, metadata often lack completeness and reliability. In this paper, we explore the feasibility of a fully automated approach for the detection of specific social events. In comparison to related approaches, we do not incorporate query-specific processing and we perform no manual adaptation of the input query. The resulting approach is applicable to arbitrary event types.
Maia Rohm, Matthias Zeppelzauer, Christian Breiteneder
ICMR2
2011 Cross-Modal Analysis of Audio-Visual Film Montage
abstract
A stylistic device frequently employed by filmmakers is the synchronous montage (composition) of audio and visual elements. Synchronous montage helps to increase tension and tempo in a scene and highlights important events in the story. Sequences with synchronous montage usually contain rich semantics which is relevant for understanding a movie. This property is currently not exploited in automated indexing, annotation, and summarization of movies. We propose a cross-modal approach that extracts sequences from a movie with synchronous audio-visual montage. Experiments confirm that the extracted sequences have high semantic relevance. Consequently, they represent a useful basis for different high-level movie abstraction tasks such as automated movie annotation and movie summarization.
Matthias Zeppelzauer, Dalibor Mitrovic, Christian Breiteneder
ICCCN1
2010 Camera Take Reconstruction
Maia Rohm, Matthias Zeppelzauer, Christian Breiteneder, Dalibor Mitrovic
MMM2
2010 A Novel Trajectory Clustering Approach for Motion Segmentation
Matthias Zeppelzauer, Maia Rohm, Dalibor Mitrovic, Christian Breiteneder
MMM1
2009 Finding the Missing Piece: Content-Based Video Comparison
abstract
The contribution of this paper consists of a framework for video comparison that allows for the analysis of different movie versions. Furthermore, a second contribution is an evaluation of state-of-the-art, local feature-based approaches for content-based video retrieval in a real world scenario. Eventually, the experimental results show the outstanding performance of a simple, edge-based descriptor within the presented framework.
Maia Rohm, Matthias Zeppelzauer, Dalibor Mitrovic, Christian Breiteneder
ISM2
2006 Discrimination and retrieval of animal sounds
abstract
Until recently few research has been performed in the area of animal sound retrieval. The authors identify state-of-the-art techniques in general purpose sound recognition by a broad survey of literature. Based on the findings, this paper gives a thorough investigation of audio features and classifiers and their applicability in the domain of animal sounds. We introduce a set of novel audio descriptors and compare their quality to other popular features. The results are encouraging and motivate further research in this domain
Dalibor Mitrovic, Matthias Zeppelzauer, Christian Breiteneder
MMM2