VLDB 2026 Research / reviewers in the wild / expert
Alexander Schindler
dblp:68/10703
· DBLP profile ↗
16ranked-venue papers
6as first author
8since 2021 · last 2024
0000-0002-4881-6741ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Taxonomap: an Interactive System for the Exploration and Explanation of Unsupervised Large-Scale News ClassificationabstractCreating analysis reports on events published in open source news data is a tedious task when done manually. Due to the large-scale nature of news data, analysts, such as government officials, often spend unnecessary resources when trying to research news data on a specific topic. In this paper, we present an interactive system for unsupervised classification of news articles in a dynamic set of hierarchical labels. By providing users with explanations in the form of highlighted words, we enable them to quickly assess the relevance of an article to a particular topic. We also provide aggregated visualisations to detect emerging events and include several quality-of-life enhancements such as a source rating mechanism and report generation. Simon Ott, Daria Liakhovets, Mina Schütz, Medina Andresel, Moritz W. Rothmund-Burgwall, Armin Vogl, Heidi Scheichenbauer, Michael Suker, Alexander Schindler |
CBMI | 9 |
| 2024 | Toolchain for Comprehensive Audio/Video Analysis Using Deep Learning Based Multimodal Approach: Use Case of Riot or Violent Context DetectionabstractIn this paper, we present a toolchain for a comprehensive audio/video analysis by leveraging deep learning based multimodal approach. To this end, different specific tasks of Speech to Text (S2T), Acoustic Scene Classification (ASC), Acoustic Event Detection (AED), Visual Object Detection (VOD), Image Captioning (IC), and Video Captioning (VC) are conducted and integrated into the toolchain. By combining individual tasks and analyzing both audio & visual data extracted from input video, the toolchain offers various audio/video-based applications: Two general applications of audio/video clustering, comprehensive audio/video summary and a specific application of riot or violent context detection. Furthermore, the toolchain presents a flexible and adaptable architecture that is effective to integrate new models for further audio/video-based applications. Lam Pham, Tin Nguyen 0007, Phat Lam, Hieu Tang, Alexander Schindler |
CBMI | 5 |
| 2024 | GerDISDETECT: A German Multilabel Dataset for Disinformation DetectionabstractDisinformation has become increasingly relevant in recent years both as a political issue and as object of research. Datasets for training machine learning models, especially for other languages than English, are sparse and the creation costly. Annotated datasets often have only binary or multiclass labels, which provide little information about the grounds and system of such classifications. We propose a novel textual dataset GerDISDETECT for German disinformation. To provide comprehensive analytical insights, a fine-grained taxonomy guided annotation scheme is required. The goal of this dataset, instead of providing a direct assessment regarding true or false, is to provide wide-ranging semantic descriptors that allow for complex interpretation as well as inferred decision-making regarding information and trustworthiness of potentially critical articles. This allows this dataset to be also used for other tasks. The dataset was collected in the first three months of 2022 and contains 39 multilabel classes with 5 top-level categories for a total of 1,890 articles: General View (3 labels), Offensive Language (11 labels), Reporting Style (15 labels), Writing Style (6 labels), and Extremism (4 labels). As a baseline, we further pre-trained a multilingual XLM-R model on around 200,000 unlabeled news articles and fine-tuned it for each category. Mina Schütz, Daniela Pisoiu, Daria Liakhovets, Alexander Schindler, Melanie Siegel |
LREC/COLING | 4 |
| 2024 | LSTM-based Deep Neural Network With A Focus on Sentence Representation for Sequential Sentence Classification in Medical Scientific AbstractsabstractThe Sequential Sentence Classification task within the domain of medical abstracts, termed as SSC, involves the categorization of sentences into pre-defined headings based on their roles in conveying critical information in the abstract.In the SSC task, sentences are sequentially related to each other.For this reason, the role of sentence embeddings is crucial for capturing both the semantic information between words in the sentence and the contextual relationship of sentences within the abstract, which then enhances the SSC system performance.In this paper, we propose a LSTM-based deep learning network with a focus on creating comprehensive sentence representation at the sentence level.To demonstrate the efficacy of the created sentence representation, a system utilizing these sentence embeddings is also developed, which consists of a Convolutional-Recurrent neural network (C-RNN) at the abstract level and a multi-layer perception network (MLP) at the segment level.Our proposed system yields highly competitive results compared to state-ofthe-art systems and further enhances the F1 scores of the baseline by 1.0%, 2.8%, and 2.6% on the benchmark datasets PudMed 200K RCT, PudMed 20K RCT and NICTA-PIBOSO, respectively.This indicates the significant impact of improving sentence representation on boosting model performance. Phat Lam, Lam Pham, Tin Nguyen 0007, Hieu Tang, Michael Seidl, Medina Andresel, Alexander Schindler |
FedCSIS | 7 |
| 2024 | Landslide Detection and Segmentation Using Remote Sensing Images and Deep Neural NetworksabstractKnowledge about historic landslide event occurrences is important for supporting disaster risk reduction strategies. Building upon findings from the 2022 Landslide4Sense competition, we propose a workflow based on a deep neural network architecture for landslide detection and segmentation from multi-source remote sensing image input. We use a U-Net trained with cross entropy loss as baseline model. We then improve this model by leveraging a wide range of deep learning techniques. In particular, we conduct feature engineering by generating new band data from the original bands, which helps to enhance the quality of the remote sensing image input. Regarding the network architecture, we replace traditional convolutional layers in the U-Net baseline by a residual-convolutional layer. We also propose an attention layer, which leverages the multi-head attention scheme. Additionally, we generate multiple output masks with three different resolutions, which creates an ensemble of three outputs in the inference process to enhance the performance. Finally, we propose a combined loss function, which leverages on focal loss and IoU loss to train the network. Our experiments on the development set of the Landslide4Sense challenge achieve an F1 score and an mIoU score of 84.07 and 76.07, respectively. Our best model setup outperforms the challenge baseline and the proposed U-Net baseline, improving the F1 score and the mIoU score by 6.8/7.4 and 10.5/8.8, respectively. Cam Le, Lam Pham, Jasmin Lampert, Matthias Schlögl, Alexander Schindler |
IGARSS | 5 |
| 2023 | Deep Learning Based Multimodal with Two-phase Training Strategy for Daily Life Video ClassificationabstractIn this paper, we present a deep learning based multimodal system for classifying daily life videos. To train the system, we propose a two-phase training strategy. In the first training phase (Phase I), we extract the audio and visual (image) data from the original video. We then train the audio data and the visual data with independent deep learning based models. After the training processes, we obtain audio embeddings and visual embeddings by extracting feature maps from the pre-trained deep learning models. In the second training phase (Phase II), we train a fusion layer to combine the audio/visual embeddings and a dense layer to classify the combined embedding into target daily scenes. Our extensive experiments, which were conducted on the benchmark dataset of DCASE (IEEE AASP Challenge on Detection and Classification of Acoustic Scenes and Events) 2021 Task 1B Development, achieved the best classification accuracy of 80.5%, 91.8%, and 95.3% with only audio data, with only visual data, both audio and visual data, respectively. The highest classification accuracy of 95.3% presents an improvement of 17.9% compared with DCASE baseline and shows very competitive to the state-of-the-art systems. Lam Pham, Trang Le, Cam Le, Dat Ngo, Axel Weissenfeld, Alexander Schindler |
CBMI | 6 |
| 2022 | An Audio-Visual Dataset and Deep Learning Frameworks for Crowded Scene ClassificationabstractIn this paper, we present the task of audio-visual scene classification (SC) where input videos are classified into one of five real-life crowded scenes: ‘Riot’, ‘Noise-Street’, ‘Firework-Event’, ‘Music-Event’, and ‘Sport-Atmosphere’. To this end, we firstly collect an audio-visual dataset (videos) of these five crowded contexts from Youtube (in-the-wild scenes). Then, a wide range of deep learning classification models are proposed to train either audio or visual input data independently. Finally, results obtained from high-performance models are fused to achieve the best accuracy score. Our experimental results indicate that audio and visual input factors independently contribute to the SC task’s performance. Notably, an ensemble of deep learning models can achieve the best accuracy of 95.7%. Lam Pham, Dat Ngo, Tho Nguyen, Phu X. Nguyen 0001, Truong Van Hoang, Alexander Schindler |
CBMI | 6 |
| 2022 | Wider or Deeper Neural Network Architecture for Acoustic Scene Classification with Mismatched Recording DevicesabstractIn this paper, we present a robust and low complexity model for Acoustic Scene Classification (ASC), the task of identifying the scene of an audio recording. We firstly construct an ASC model in which a novel inception-residual-based network architecture is proposed to deal with the issue of mismatched recording devices. To further improve the model performance but still satisfy the low footprint, we apply two techniques of ensemble of multiple spectrograms and model compression to the proposed ASC model. By conducting extensive experiments on the benchmark DCASE 2020 Task 1A Development dataset, we achieve the best model performing an accuracy of 71.3% and a low complexity of 0.5 Million (M) trainable parameters, which is very competitive to the state-of-the-art systems and potential for real-life applications on edge devices. Lam Pham, Dat Ngo, Hieu Tang, Son Phan, Alexander Schindler |
MMAsia | 6 |
| 2020 | Multi-modal video forensic platform for investigating post-terrorist attack scenariosabstractThe forensic investigation of a terrorist attack poses a significant challenge to the investigative authorities, as often several thousand hours of video footage must be viewed. Large scale Video Analytic Platforms (VAP) assist law enforcement agencies (LEA) in identifying suspects and securing evidence. Current platforms focus primarily on the integration of different computer vision methods and thus are restricted to a single modality. We present a video analytic platform that integrates visual and audio analytic modules and fuses information from surveillance cameras and video uploads from eyewitnesses. Videos are analyzed according their acoustic and visual content. Specifically, Audio Event Detection is applied to index the content according to attack-specific acoustic concepts. Audio similarity search is utilized to identify similar video sequences recorded from different perspectives. Visual object detection and tracking are used to index the content according to relevant concepts. Innovative user-interface concepts are introduced to harness the full potential of the heterogeneous results of the analytical modules, allowing investigators to more quickly follow-up on leads and eyewitness reports. Alexander Schindler, Andrew Lindley, Anahid N. Jalali, Martin Boyer, Sergiu Gordea, Ross King |
MMSys | 1 |
| 2019 | Multi-Task Music Representation Learning from Multi-Label EmbeddingsabstractThis paper presents a novel approach to music representation learning. Triplet loss based networks have become popular for representation learning in various multimedia retrieval domains. Yet, one of the most crucial parts of this approach is the appropriate selection of triplets, which is indispensable, considering that the number of possible triplets grows cubically. We present an approach to harness multi-tag annotations for triplet selection, by using Latent Semantic Indexing to project the tags onto a high-dimensional space. From this we estimate tag-relatedness to select hard triplets. The approach is evaluated in a multi-task scenario for which we introduce four large multi-tag annotations for the Million Song Dataset for the music properties genres, styles, moods, and themes. Alexander Schindler, Peter Knees |
CBMI | 1 |
| 2019 | Large Scale Audio-Visual Video Analytics Platform for Forensic Investigations of Terroristic Attacks
Alexander Schindler, Martin Boyer, Andrew Lindley, David Schreiber, Thomas Philipp |
MMM (2) | 1 |
| 2019 | On the Unsolved Problem of Shot Boundary Detection for Music Videos
Alexander Schindler, Andreas Rauber |
MMM (1) | 1 |
| 2017 | Harnessing Music-Related Visual Stereotypes for Music Information RetrievalabstractOver decades, music labels have shaped easily identifiable genres to improve recognition value and subsequently market sales of new music acts. Referring to print magazines and later to music television as important distribution channels, the visual representation thus played and still plays a significant role in music marketing. Visual stereotypes developed over decades that enable us to quickly identify referenced music only by sight without listening. Despite the richness of music-related visual information provided by music videos and album covers as well as T-shirts, advertisements, and magazines, research towards harnessing this information to advance existing or approach new problems of music retrieval or recommendation is scarce or missing. In this article, we present our research on visual music computing that aims to extract stereotypical music-related visual information from music videos. To provide comprehensive and reproducible results, we present the Music Video Dataset, a thoroughly assembled suite of datasets with dedicated evaluation tasks that are aligned to current Music Information Retrieval tasks. Based on this dataset, we provide evaluations of conventional low-level image processing and affect-related features to provide an overview of the expressiveness of fundamental visual properties such as color, illumination, and contrasts. Further, we introduce a high-level approach based on visual concept detection to facilitate visual stereotypes. This approach decomposes the semantic content of music video frames into concrete concepts such as vehicles, tools, and so on, defined in a wide visual vocabulary. Concepts are detected using convolutional neural networks and their frequency distributions as semantic descriptions for a music video. Evaluations showed that these descriptions show good performance in predicting the music genre of a video and even outperform audio-content descriptors on cross-genre thematic tags. Further, highly significant performance improvements were observed by augmenting audio-based approaches through the introduced visual approach. Alexander Schindler, Andreas Rauber |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2015 | An Audio-Visual Approach to Music Genre Classification through Affective Color Features
Alexander Schindler, Andreas Rauber |
ECIR | 1 |
| 2013 | Duplicate detection approaches for quality assurance of document image collectionsabstractThis paper presents an evaluation of different methods for automatic duplicate detection in digitized collections. These approaches are meant to support quality assurance and decision making for long term preservation of digital content in libraries and archives. In this paper we demonstrate advantages and drawbacks of different approaches. Our goal is to select the most efficient method which satisfies the digital preservation requirements for duplicate detection in digital document image collections. Workflows of different complexity were designed in order to demonstrate possible duplicate detection approaches. Assessment of individual approaches is based on workflow simplicity, detection accuracy and acceptable performance, since image processing methods typically require significant computation. Applied image processing methods create expert knowledge that facilitates decision making for long term preservation. We employ AI technologies like expert rules and clustering for inferring explicit knowledge on the content of the digital collection. A statistical analysis of the aggregated information and the qualitative analysis of the aggregated knowledge are presented in the evaluation part of the paper. Roman Graf, Reinhold Huber-Mörk, Alexander Schindler, Sven Schlarb |
MEDES | 3 |
| 2012 | Quality Assurance for Document Image Collections in Digital Preservation
Reinhold Huber-Mörk, Alexander Schindler |
ACIVS | 2 |