EDBT 2026 Demo / reviewers in the wild / expert
Gerhard Widmer
dblp:w/GerhardWidmer
· DBLP profile ↗
94ranked-venue papers
24as first author
17since 2021 · last 2026
0000-0003-3531-1282ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 54 · 20 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 5 first-author · 11 since 2021Databases, data management, data science and information retrieval · 21 · 5 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 3Theory of computation · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Sound Event Detection With Boundary-Aware Optimization and InferenceabstractTemporal detection problems appear in many fields including time-series estimation, activity recognition and sound event detection (SED). In this work, we propose a new approach to temporal event modeling by explicitly modeling event onsets and offsets, and by introducing boundary-aware optimization and inference strategies that substantially enhance temporal event detection. The presented methodology incorporates new temporal modeling layers—Recurrent Event Detection (RED) and Event Proposal Network (EPN)—which, together with tailored loss functions, enable more effective and precise temporal event detection. We evaluate the proposed method in the SED domain using a subset of the temporally-strongly annotated portion of AudioSet. Experimental results show that our approach not only outperforms traditional frame-wise SED models with state-of-the-art post-processing, but also removes the need for post-processing hyperparameter tuning, and scales to achieve new state-of-the-art performance across all AudioSet Strong classes. Florian Schmid, Chi Ian Tang, Sanjeel Parekh, Vamsi K. Ithapu, Juan Azcarreta, Giacomo Ferroni, Yijun Qian, Arnoldas Jasonas, Cosmin Frateanu, Camilla Clark, Gerhard Widmer, Cagdas Bilen |
IEEE Signal Process. Lett. | 11 |
| 2025 | Estimating Musical Surprisal in AudioabstractIn modeling musical surprisal expectancy with computational methods, it has been proposed to use the information content (IC) of one-step predictions from an autoregressive model as a proxy for surprisal in symbolic music. With an appropriately chosen model, the IC of musical events has been shown to correlate with human perception of surprise and complexity aspects, including tonal and rhythmic complexity. This work investigates whether an analogous methodology can be applied to music audio. We train an autoregressive Transformer model to predict compressed latent audio representations of a pretrained autoencoder network. We verify learning effects by estimating the decrease in IC with repetitions. We investigate the mean IC of musical segment types (e.g., A or B) and find that segment types appearing later in a piece have a higher IC than earlier ones on average. We investigate the IC’s relation to audio and musical features and find it correlated with timbral variations and loudness and, to a lesser extent, dissonance, rhythmic complexity, and onset density related to audio and musical features. Finally, we investigate if the IC can predict EEG responses to songs and thus model humans’ surprisal in music. We provide code for our method on github.com/sonycslparis/audioic. Mathias Rose Bjare, Giorgia Cantisani, Stefan Lattner, Gerhard Widmer |
ICASSP | 4 |
| 2025 | Effective Pre-Training of Audio Transformers for Sound Event DetectionabstractWe propose a pre-training pipeline for audio spectrogram transformers for frame-level sound event detection tasks. On top of common pre-training steps, we add a meticulously designed training routine on AudioSet frame-level annotations. This includes a balanced sampler, aggressive data augmentation, and ensemble knowledge distillation. For five transformers, we obtain a substantial performance improvement over previously available checkpoints both on AudioSet frame-level predictions and on frame-level sound event detection downstream tasks, confirming our pipeline’s effectiveness. We publish the resulting checkpoints that researchers can directly fine-tune to build high-performance models for sound event detection tasks. Florian Schmid, Tobias Morocutti, Francesco Foscarin, Jan Schlüter, Paul Primus, Gerhard Widmer |
ICASSP | 6 |
| 2024 | Perception-Inspired Graph Convolution for Music Understanding Tasks
Emmanouil Karystinaios, Francesco Foscarin, Gerhard Widmer |
IJCAI | 3 |
| 2024 | Rethinking data augmentation for adversarial robustness
Hamid Eghbalzadeh, Werner Zellinger, Maura Pintor, Kathrin Grosse, Khaled Koutini, Bernhard Moser 0001, Battista Biggio, Gerhard Widmer |
Inf. Sci. | 8 |
| 2024 | Dynamic Convolutional Neural Networks as Efficient Pre-Trained Audio ModelsabstractThe introduction of large-scale audio datasets, such as AudioSet, paved the way for Transformers to conquer the audio domain and replace CNNs as the state-of-the-art neural network architecture for many tasks. Audio Spectrogram Transformers are excellent at exploiting large datasets, creating powerful pre-trained models that surpass CNNs when fine-tuned on downstream tasks. However, current popular Audio Spectrogram Transformers are demanding in terms of computational complexity compared to CNNs. Recently, we have shown that, by employing Transformer-to-CNN Knowledge Distillation, efficient CNNs can catch up with and even outperform Transformers on large datasets. In this work, we extend this line of research and increase the capacity of efficient CNNs by introducing dynamic CNN blocks constructed of dynamic convolutions, a dynamic ReLU activation function, and Coordinate Attention. We show that these dynamic CNNs outperform traditional efficient CNNs, such as MobileNets, in terms of the performance–complexity trade-off at the task of audio tagging on the large-scale AudioSet. Our experiments further indicate that the proposed dynamic CNNs achieve competitive performance with Transformer-based models for end-to-end fine-tuning on downstream tasks while being much more computationally efficient. Florian Schmid, Khaled Koutini, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Efficient Large-Scale Audio Tagging Via Transformer-to-CNN Knowledge DistillationabstractAudio Spectrogram Transformer models rule the field of Audio Tagging, outrunning previously dominating Convolutional Neural Networks (CNNs). Their superiority is based on the ability to scale up and exploit large-scale datasets such as AudioSet. However, Transformers are demanding in terms of model size and computational requirements compared to CNNs. We propose a training procedure for efficient CNNs based on offline Knowledge Distillation (KD) from high-performing yet complex transformers. The proposed training schema and the efficient CNN design based on MobileNetV3 results in models outperforming previous solutions in terms of parameter and computational efficiency and prediction performance. We provide models of different complexity levels, scaling from low-complexity models up to a new state-of-the-art performance of .483 mAP on AudioSet.1 Florian Schmid, Khaled Koutini, Gerhard Widmer |
ICASSP | 3 |
| 2023 | The ACCompanion: Combining Reactivity, Robustness, and Musical Expressivity in an Automatic Piano AccompanistabstractThis paper introduces the ACCompanion, an expressive accompaniment system. Similarly to a musician who accompanies a soloist playing a given musical piece, our system can produce a human-like rendition of the accompaniment part that follows the soloist's choices in terms of tempo, dynamics, and articulation. The ACCompanion works in the symbolic domain, i.e., it needs a musical instrument capable of producing and playing MIDI data, with explicitly encoded onset, offset, and pitch for each played note. We describe the components that go into such a system, from real-time score following and prediction to expressive performance generation and online adaptation to the expressive choices of the human player. Based on our experience with repeated live demonstrations in front of various audiences, we offer an analysis of the challenges of combining these components into a system that is highly reactive and precise, while still a reliable musical partner, robust to possible performance errors and responsive to expressive variations. Carlos Eduardo Cancino-Chacón, Silvan Peter, Patricia Hu, Emmanouil Karystinaios, Florian Henkel, Francesco Foscarin, Gerhard Widmer |
IJCAI | 7 |
| 2023 | Musical Voice Separation as Link Prediction: Modeling a Musical Perception Task as a Multi-Trajectory Tracking ProblemabstractThis paper targets the perceptual task of separating the different interacting voices, i.e., monophonic melodic streams, in a polyphonic musical piece. We target symbolic music, where notes are explicitly encoded, and model this task as a Multi-Trajectory Tracking (MTT) problem from discrete observations, i.e., notes in a pitch-time space. Our approach builds a graph from a musical piece, by creating one node for every note, and separates the melodic trajectories by predicting a link between two notes if they are consecutive in the same voice/stream. This kind of local, greedy prediction is made possible by node embeddings created by a heterogeneous graph neural network that can capture inter- and intra-trajectory information. Furthermore, we propose a new regularization loss that encourages the output to respect the MTT premise of at most one incoming and one outgoing link for every node, favoring monophonic (voice) trajectories; this loss function might also be useful in other general MTT scenarios. Our approach does not use domain-specific heuristics, is scalable to longer sequences and a higher number of voices, and can handle complex cases such as voice inversions and overlaps. We reach new state-of-the-art results for the voice separation task on classical music of different styles. All code, data, and pretrained models are available on https://github.com/manoskary/vocsep_ijcai2023 Emmanouil Karystinaios, Francesco Foscarin, Gerhard Widmer |
IJCAI | 3 |
| 2023 | Discrete Diffusion Probabilistic Models for Symbolic Music GenerationabstractDenoising Diffusion Probabilistic Models (DDPMs) have made great strides in generating high-quality samples in both discrete and continuous domains. However, Discrete DDPMs (D3PMs) have yet to be applied to the domain of Symbolic Music. This work presents the direct generation of Polyphonic Symbolic Music using D3PMs. Our model exhibits state-of-the-art sample quality, according to current quantitative evaluation metrics, and allows for flexible infilling at the note level. We further show, that our models are accessible to post-hoc classifier guidance, widening the scope of possible applications. However, we also cast a critical view on quantitative evaluation of music sample quality via statistical metrics, and present a simple algorithm that can confound our metrics with completely spurious, non-musical samples. Matthias Plasser, Silvan Peter, Gerhard Widmer |
IJCAI | 3 |
| 2023 | Self-Supervised Contrastive Learning for Robust Audio-Sheet Music Retrieval SystemsabstractLinking sheet music images to audio recordings remains a key problem for the development of efficient cross-modal music retrieval systems. One of the fundamental approaches toward this task is to learn a cross-modal embedding space via deep neural networks that is able to connect short snippets of audio and sheet music. However, the scarcity of annotated data from real musical content affects the capability of such methods to generalize to real retrieval scenarios. In this work, we investigate whether we can mitigate this limitation with self-supervised contrastive learning, by exposing a network to a large amount of real music data as a pre-training step, by contrasting randomly augmented views of snippets of both modalities, namely audio and sheet images. Through a number of experiments on synthetic and real piano data, we show that pretrained models are able to retrieve snippets with better precision in all scenarios and pre-training configurations. Encouraged by these results, we employ the snippet embeddings in the higher-level task of cross-modal piece identification and conduct more experiments on several retrieval configurations. In this task, we observe that the retrieval quality improves from 30% up to 100% when real music data is present. We then conclude by arguing for the potential of self-supervised contrastive learning for alleviating the annotated data scarcity in multi-modal music retrieval models. Code and trained models are accessible at https://github.com/luisfvc/ucasr. Tobias Washüttl, Gerhard Widmer |
MMSys | 3 |
| 2023 | Constructing adversarial examples to investigate the plausibility of explanations in deep audio and image classifiersabstractAbstract Given the rise of deep learning and its inherent black-box nature, the desire to interpret these systems and explain their behaviour became increasingly more prominent. The main idea of so-called explainers is to identify which features of particular samples have the most influence on a classifier’s prediction, and present them as explanations. Evaluating explainers, however, is difficult, due to reasons such as a lack of ground truth. In this work, we construct adversarial examples to check the plausibility of explanations, perturbing input deliberately to change a classifier’s prediction. This allows us to investigate whether explainers are able to detect these perturbed regions as the parts of an input that strongly influence a particular classification. Our results from the audio and image domain suggest that the investigated explainers often fail to identify the input regions most relevant for a prediction; hence, it remains questionable whether explanations are useful or potentially misleading. Katharina Hoedt, Verena Praher, Arthur Flexer, Gerhard Widmer |
Neural Comput. Appl. | 4 |
| 2022 | Efficient Training of Audio Transformers with PatchoutabstractThe great success of transformer-based models in natural language processing (NLP) has led to various attempts at adapting these architectures to other domains such as vision and audio. Recent work has shown that transformers can outperform Convolutional Neural Networks (CNNs) on vision and audio tasks. However, one of the main shortcomings of transformer models, compared to the well-established CNNs, is the computational complexity. In transformers, the compute and memory complexity is known to grow quadratically with the input length. Therefore, there has been extensive work on optimizing transformers, but often at the cost of degrading predictive performance. In this work, we propose a novel method to optimize and regularize transformers on audio spectrograms. Our proposed models achieve a new state-of-the-art performance on Audioset and can be trained on a single consumer-grade GPU. Furthermore, we propose a transformer model that outperforms CNNs in terms of both performance and training speed. Source code: https://github.com/kkoutini/PaSST Khaled Koutini, Jan Schlüter, Hamid Eghbalzadeh, Gerhard Widmer |
INTERSPEECH | 4 |
| 2021 | LEMONS: Listenable Explanations for Music recOmmeNder Systems
Alessandro B. Melchiorre, Verena Praher, Markus Schedl, Gerhard Widmer |
ECIR (2) | 4 |
| 2021 | Con Espressione! AI, Machine Learning, and Musical Expressivity
Gerhard Widmer |
ICAART (1) | 1 |
| 2021 | Towards Explaining Expressive Qualities in Piano Recordings: Transfer of Explanatory Features Via Acoustic Domain AdaptationabstractEmotion and expressivity in music have been topics of considerable interest in the field of music information retrieval. In recent years, mid-level perceptual features have been suggested as means to explain computational predictions of musical emotion. We find that the diversity of musical styles and genres in the available dataset for learning these features is not sufficient for models to generalise well to specialised acoustic domains such as solo piano music. In this work, we show that by utilising unsupervised domain adaptation together with receptive-field regularised deep neural networks, it is possible to significantly improve generalisation to this domain. Additionally, we demonstrate that our domain-adapted models can better predict and explain expressive qualities in classical piano performances, as perceived and described by human listeners. Shreyan Chowdhury, Gerhard Widmer |
ICASSP | 2 |
| 2021 | Receptive Field Regularization Techniques for Audio Classification and Tagging With Deep Convolutional Neural NetworksabstractIn this paper, we study the performance of variants of well-known Convolutional Neural Network (CNN) architectures on different audio tasks. We show that tuning the Receptive Field (RF) of CNNs is crucial to their generalization. An insufficient RF limits the CNN's ability to fit the training data. In contrast, CNNs with an excessive RF tend to over-fit the training data and fail to generalize to unseen testing data. As state-of-the-art CNN architectures - in computer vision and other domains - tend to go deeper in terms of number of layers, their RF size increases and therefore they degrade in performance in several audio classification and tagging tasks. We study well-known CNN architectures and how their building blocks affect their receptive field. We propose several systematic approaches to control the RF of CNNs and systematically test the resulting architectures on different audio classification and tagging tasks and datasets. The experiments show that regularizing the RF of CNNs using our proposed approaches can drastically improve the generalization of models, out-performing complex architectures and pre-trained models on larger datasets. The proposed CNNs achieve state-of-the-art results in multiple tasks, from acoustic scene classification to emotion and theme detection in music to instrument recognition, as demonstrated by top ranks in several pertinent challenges (DCASE, MediaEval). Khaled Koutini, Hamid Eghbalzadeh, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Mixture Density Generative Adversarial NetworksabstractGenerative Adversarial Networks have a surprising ability to generate sharp and realistic images, but they are known to suffer from the so-called mode collapse problem. In this paper, we propose a new GAN variant called Mixture Density GAN that overcomes this problem by encouraging the Discriminator to form clusters in its embedding space, which in turn leads the Generator to exploit these and discover different modes in the data. This is achieved by positioning Gaussian density functions in the corners of a simplex, using the resulting Gaussian mixture as a likelihood function over discriminator embeddings, and formulating an objective function for GAN training that is based on these likelihoods. We show how formation of these clusters changes the probability landscape of the discriminator and improves the mode discovery of the GAN. We also show that the optimum of our training objective is attained if and only if the generated and the real distribution match exactly. We support our theoretical results with empirical evaluations on three mode discovery benchmark datasets (Stacked-MNIST, Ring of Gaussians and Grid of Gaussians), and four image datasets (CIFAR-10, CelebA, MNIST, and Fashion-MNIST). Furthermore, we demonstrate (1) the ability to avoid mode collapse and discover all the modes and (2) superior quality of the generated images (as measured by the Fréchet Inception Distance (FID)), achieving the lowest FID compared to all baselines. Hamid Eghbalzadeh, Werner Zellinger, Gerhard Widmer |
CVPR | 3 |
| 2019 | Deep Polyphonic ADSR Piano Note TranscriptionabstractWe investigate a late-fusion approach to piano transcription, combined with a strong temporal prior in the form of a handcrafted Hidden Markov Model (HMM). The network architecture under consideration is compact in terms of its number of parameters and easy to train with gradient descent. The network outputs are fused over time in the final stage to obtain note segmentations, with an HMM whose transition probabilities are chosen based on a model of attack, decay, sustain, release (ADSR) envelopes, commonly used for sound synthesis. The note segments are then subject to a final binary decision rule to reject too weak note segment hypotheses. We obtain state-of-the-art results on the MAPS dataset, and are able to outperform other approaches by a large margin, when predicting complete note regions from onsets to offsets. Rainer Kelz, Sebastian Böck, Gerhard Widmer |
ICASSP | 3 |
| 2019 | Feature-combination hybrid recommender systems for automated music playlist continuationabstractMusic recommender systems have become a key technology to support the interaction of users with the increasingly larger music catalogs of on-line music streaming services, on-line music shops, and personal devices. An important task in music recommender systems is the automated continuation of music playlists, that enables the recommendation of music streams adapting to given (possibly short) listening sessions. Previous works have shown that applying collaborative filtering to collections of curated music playlists reveals underlying playlist-song co-occurrence patterns that are useful to predict playlist continuations. However, most music collections exhibit a pronounced long-tailed distribution. The majority of songs occur only in few playlists and, as a consequence, they are poorly represented by collaborative filtering. We introduce two feature-combination hybrid recommender systems that extend collaborative filtering by integrating the collaborative information encoded in curated music playlists with any type of song feature vector representation. We conduct off-line experiments to assess the performance of the proposed systems to recover withheld playlist continuations, and we compare them to competitive pure and hybrid collaborative filtering baselines. The results of the experiments indicate that the introduced feature-combination hybrid recommender systems can more accurately predict fitting playlist continuations as a result of their improved representation of songs occurring in few playlists. Andreu Vall, Matthias Dorfer, Hamid Eghbalzadeh, Markus Schedl, Keki Burjorjee, Gerhard Widmer |
User Model. User Adapt. Interact. | 6 |
| 2018 | Investigating Label Noise Sensitivity of Convolutional Neural Networks for Fine Grained Audio Signal LabellingabstractWe measure the effect of small amounts of systematic and random label noise caused by slightly misaligned ground truth labels in a fine grained audio signal labeling task. The task we choose to demonstrate these effects on is also known as framewise polyphonic transcription or note quantized multi-fO estimation, and transforms a monaural audio signal into a sequence of note indicator labels. It will be shown that even slight misalignments have clearly apparent effects, demonstrating a great sensitivity of convolutional neural networks to label noise. The implications are clear: when using convolutional neural networks for fine grained audio signal labeling tasks, great care has to be taken to ensure that the annotations have precise timing, and are free from systematic or random error as much as possible - even small misalignments will have a noticeable impact. Rainer Kelz, Gerhard Widmer |
ICASSP | 2 |
| 2018 | A Large-Scale Study of Language Models for Chord PredictionabstractWe conduct a large-scale study of language models for chord prediction. Specifically, we compare N-gram models to various flavours of recurrent neural networks on a comprehensive dataset comprising all publicly available datasets of annotated chords known to us. This large amount of data allows us to systematically explore hyperparameter settings for the recurrent neural networks-a crucial step in achieving good results with this model class. Our results show not only a quantitative difference between the models, but also a qualitative one: in contrast to static N-gram models, certain RNN configurations adapt to the songs at test time. This finding constitutes a further step towards the development of chord recognition systems that are more aware of local musical context than what was previously possible. Filip Korzeniowski, David R. W. Sears, Gerhard Widmer |
ICASSP | 3 |
| 2018 | Machine Learning Approaches to Hybrid Music Recommender Systems
Andreu Vall, Gerhard Widmer |
ECML/PKDD (3) | 2 |
| 2018 | Online, Loudness-Invariant Vocal Detection in Mixed Music SignalsabstractSinging voice detection, also referred to as vocal detection (VD), aims at automatically identifying the regions in a music recording where at least one person sings. It is highly challenging due to the timbral and expressive richness of the human singing voice, as well as the practically endless variety of interfering instrumental accompaniment. Additionally, certain instruments have an inherent risk of being misclassified as vocals due to similarities of the sound production system. In this paper, we present a machine learning approach that is based on our previous work for VD, which is specifically designed to deal with those challenging conditions. The contribution of this paper is threefold: First, we present a new method for VD that passes a compact set of features to a long short-term memory recurrent neural network classifier that obtains state-of-the-art results. Second, we thoroughly evaluate the proposed method along with related approaches to really probe the weaknesses of the methods. In order to allow for such a thorough evaluation, we make a curated collection of datasets available to the research community. Finally, we focus on a specific problem that was not obvious and had not been discussed in the literature so far. The reason for this is precisely because limited evaluations had not revealed this as a problem: the lack of loudness invariance. We will discuss the implications of utilizing loudness-related features and show that our method successfully deals with this problem due to the specific set of features it uses. Bernhard Lehner, Jan Schlüter, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | A Review of Automatic Drum TranscriptionabstractIn Western popular music, drums and percussion are an important means to emphasize and shape the rhythm, often defining the musical style. If computers were able to analyze the drum part in recorded music, it would enable a variety of rhythm-related music processing tasks. Especially the detection and classification of drum sound events by computational methods is considered to be an important and challenging research problem in the broader field of music information retrieval. Over the last two decades, several authors have attempted to tackle this problem under the umbrella term automatic drum transcription (ADT). This paper presents a comprehensive review of ADT research, including a thorough discussion of the task-specific challenges, categorization of existing techniques, and evaluation of several state-of-the-art systems. To provide more insights on the practice of ADT systems, we focus on two families of ADT techniques, namely methods based on non-negative matrix factorization and recurrent neural networks. We explain the methods’ technical details and drum-specific variations and evaluate these approaches on publicly available data sets with a consistent experimental setup. Finally, the open issues and underexplored areas in ADT research are identified and discussed, providing future directions in this field. Chih-Wei Wu, Christian Dittmar, Carl Southall, Richard Vogl, Gerhard Widmer, Jason Hockman, Meinard Müller, Alexander Lerch 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | An evaluation of linear and non-linear models of expressive dynamics in classical piano and symphonic musicabstractExpressive interpretation forms an important but complex aspect of music, particularly in Western classical music. Modeling the relation between musical expression and structural aspects of the score being performed is an ongoing line of research. Prior work has shown that some simple numerical descriptors of the score (capturing dynamics annotations and pitch) are effective for predicting expressive dynamics in classical piano performances. Nevertheless, the features have only been tested in a very simple linear regression model. In this work, we explore the potential of non-linear and temporal modeling of expressive dynamics. Using a set of descriptors that capture different types of structure in the musical score, we compare linear and different non-linear models in a large-scale evaluation on three different corpora, involving both piano and orchestral music. To the best of our knowledge, this is the first study where models of musical expression are evaluated on both types of music. We show that, in addition to being more accurate, non-linear models describe interactions between numerical descriptors that linear models do not. Carlos Eduardo Cancino-Chacón, Thassilo Gadermaier, Gerhard Widmer, Maarten Grachten |
Mach. Learn. | 3 |
| 2017 | Getting Closer to the Essence of Music: The Con Espressione ManifestoabstractThis text offers a personal and very subjective view on the current situation of Music Information Research (MIR). Motivated by the desire to build systems with a somewhat deeper understanding of music than the ones we currently have, I try to sketch a number of challenges for the next decade of MIR research, grouped around six simple truths about music that are probably generally agreed on but often ignored in everyday research. Gerhard Widmer |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2016 | madmom: A New Python Audio and Music Signal Processing LibraryabstractIn this paper, we present madmom, an open-source audio processing and music information retrieval (MIR) library written in Python. madmom features a concise, NumPy-compatible, object oriented design with simple calling conventions and sensible default values for all parameters, which facilitates fast prototyping of MIR applications. Prototypes can be seamlessly converted into callable processing pipelines through madmom's concept of Processors, callable objects that run transparently on multiple cores. Processors can also be serialised, saved, and re-run to allow results to be easily reproduced anywhere. Apart from low-level audio processing, madmom puts emphasis on musically meaningful high-level features. Many of these incorporate machine learning techniques and madmom provides a module that implements some methods commonly used in MIR such as hidden Markov models and neural networks. Additionally, madmom comes with several state-of-the-art MIR algorithms for onset detection, beat, downbeat and meter tracking, tempo estimation, and chord recognition. These can easily be incorporated into bigger MIR systems or run as stand-alone programs. Sebastian Böck, Filip Korzeniowski, Jan Schlüter, Florian Krebs, Gerhard Widmer |
ACM Multimedia | 5 |
| 2016 | Robust Quad-Based Audio FingerprintingabstractWe propose an audio fingerprinting method that adapts findings from the field of blind astrometry to define simple, efficiently representable characteristic feature combinations called quads. Based on these, an audio identification algorithm is described that is robust to noise and severe time-frequency scale distortions and accurately identifies the underlying scale transform factors. The low number and compact representation of content features allows for efficient application of exact fixed-radius near-neighbor search methods for fingerprint matching in large audio collections. We demonstrate the practicability of the method on a collection of 100,000 songs, analyze its performance for a diverse set of noise as well as severe speed, tempo and pitch scale modifications, and identify a number of advantages of our method over two state-of-the-art distortion-robust audio identification algorithms. Reinhard Sonnleitner, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Artificial Intelligence in the Concertgebouw
Andreas Arzt, Harald Frostel, Thassilo Gadermaier, Martin Gasser, Maarten Grachten, Gerhard Widmer |
IJCAI | 6 |
| 2015 | Improving voice activity detection in moviesabstractVoice Activity Detection in movies is a non-trivial and challenging task. The different emotional states of the speakers, as well as the variety of soundscapes and noises contribute to the complexity of the task. In this paper, we propose a set of lightweight features that are specifically designed to perform under such conditions, while at the same time preventing confusions of singing voice with speech. For evaluation, we use four fulllength movies, previously unseen to the system and painstakingly annotated. We compare our detector to a state-of-the-art reference system. The new approach performs better, yielding just about half the Equal Error Rate (EER). Furthermore, since the ground truth annotation task is extremely tedious, and to help with advancing in this topic, we release the annotations of all four movies to the research community. Index Terms: Voice Activity Detection, Speech Detection Bernhard Lehner, Gerhard Widmer, Reinhard Sonnleitner |
INTERSPEECH | 2 |
| 2015 | Inferring Metrical Structure in Music Using Particle FiltersabstractIn this paper, we propose a new state-of-the-art particle filter (PF) system to infer the metrical structure of musical audio signals. The new inference method is designed to overcome the problem of PFs in multi-modal probability distributions, which arise due to tempo and phase ambiguities in musical rhythm representations. We compare the new method with a hidden Markov model (HMM) system and several other PF schemes in terms of performance, speed and scalability on several audio datasets. We demonstrate that using the proposed system the computational complexity can be reduced drastically in comparison to the HMM while maintaining the same order of beat tracking accuracy. Therefore, for the first time, the proposed system allows fast meter inference in a high-dimensional state space, spanned by the three components of tempo, type of rhythm, and position in a metric cycle. Florian Krebs, Andre Holzapfel, A. Taylan Cemgil, Gerhard Widmer |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2014 | Quad-Based Audio Fingerprinting Robust to Time and Frequency Scaling
Reinhard Sonnleitner, Gerhard Widmer |
DAFx | 2 |
| 2014 | The Piano Music CompanionabstractWe present a system that we call ‘The Piano Music Companion’ and that is able to follow and understand (at least to some extent) a live piano performance. Within a few seconds this system can identify the piece that is being played, and the position within the piece. It then tracks the progress of the performer over time via a robust score following algorithm. The companion is useful in multiple ways, e.g., it can be used for piece identification, music visualisation, during piano rehearsal and for automatic page turning. Andreas Arzt, Sebastian Böck, Sebastian Flossmann, Harald Frostel, Martin Gasser, Cynthia C. S. Liem, Gerhard Widmer |
ECAI | 7 |
| 2014 | On the reduction of false positives in singing voice detectionabstractMotivated by the observation that one of the biggest problems in automatic singing voice detection is the confusion of vocals with other pitch-continuous and pitch-varying instruments, we propose a set of three new audio features designed to reduce the amount of false vocal detections. This is borne out in comparative experiments with three different musical corpora. The resulting singing voice detector appears to be at least on par with more complex state-of-the-art methods. New features and classifier are very light-weight and in principle suitable for on-line use. Bernhard Lehner, Gerhard Widmer, Reinhard Sonnleitner |
ICASSP | 2 |
| 2012 | Local and global scaling reduce hubs in space
Dominik Schnitzer, Arthur Flexer, Markus Schedl, Gerhard Widmer |
J. Mach. Learn. Res. | 4 |
| 2012 | A fast audio similarity retrieval method for millions of music tracks
Dominik Schnitzer, Arthur Flexer, Gerhard Widmer |
Multim. Tools Appl. | 3 |
| 2011 | A music information system automatically generated via Web content mining techniques
Markus Schedl, Gerhard Widmer, Peter Knees, Tim Pohle |
Inf. Process. Manag. | 2 |
| 2011 | Exploring the music similarity space on the webabstractThis article comprehensively addresses the problem of similarity measurement between music artists via text-based features extracted from Web pages. To this end, we present a thorough evaluation of different term-weighting strategies, normalization methods, aggregation functions, and similarity measurement techniques. In large-scale genre classification experiments carried out on real-world artist collections, we analyze several thousand combinations of settings/parameters that influence the similarity calculation process, and investigate in which way they impact the quality of the similarity estimates. Accurate similarity measures for music are vital for many applications, such as automated playlist generation, music recommender systems, music information systems, or intelligent user interfaces to access music collections by means beyond text-based browsing. Therefore, by exhaustively analyzing the potential of text-based features derived from artist-related Web pages, this article constitutes an important contribution to context-based music information research. Markus Schedl, Tim Pohle, Peter Knees, Gerhard Widmer |
ACM Trans. Inf. Syst. | 4 |
| 2010 | Three web-based heuristics to determine a person's or institution's country of originabstractWe propose three heuristics to determine the country of origin of a person or institution via text-based IE from the Web. We evaluate all methods on a collection of music artists and bands, and show that some heuristics outperform earlier work on the topic by terms of coverage, while retaining similar precision levels. We further investigate an extension using country-specific synonym lists. Markus Schedl, Klaus Seyerlehner, Dominik Schnitzer, Gerhard Widmer, Cornelia Schiketanz |
SIGIR | 4 |
| 2009 | Dealing with Music in Intelligent Ways
Gerhard Widmer |
ISMIS | 1 |
| 2009 | On the limitations of browsing top-N recommender systemsabstractTo exploit the enormous potential of niche products, modern information systems must support users in exploring digital libraries and online catalogs. A straight-forward way of doing so is to support browsing the available items, which is in general realized by presenting a user the top-N recommendations for each item. However, recent research indicates that most of the niche products reside in the so-called Long Tail, and simple collaborative filtering-based recommender systems alone do not allow to explore these niche products. In this paper we show that it is not only a popularity problem related to the collaborative filtering approach that makes a portion of the elements of a digital library inaccessible via browsing, but also a consequence of the top N-recommendation approach itself. Klaus Seyerlehner, Arthur Flexer, Gerhard Widmer |
RecSys | 3 |
| 2008 | Automatic Page Turning for Musicians via Real-Time Machine ListeningabstractWe present a system that automatically turns the pages of the music score for musicians during a performance. It is based on a new algorithm for following an incoming audio stream in real time and aligning it to a music score (in the form of a synthesised audio file). Precision and robustness of the algorithm are quantified in systematic experiments, and a demonstration using an actual page turning machine built by an Austrian company is described. Andreas Arzt, Gerhard Widmer, Simon Dixon |
ECAI | 2 |
| 2008 | Towards an Automatically Generated Music Information System Via Web Content Mining
Markus Schedl, Peter Knees, Tim Pohle, Gerhard Widmer |
ECIR | 4 |
| 2008 | Sound/tracks: real-time synaesthetic sonification of train journeysabstractTravelling on a train and looking out of the window at the moving scenery reveals a composition of "visual music" with its own tempo and rhythm, its own colours and harmonies. The project sound/tracks aims at capturing these visual impressions and translates them into a musical composition in real-time - producing an immediate and unique soundtrack to the train journey based on the passing landscape. To this end, the outside impressions are captured with a camera and translated into instantaneously played back piano music. The immediately added sound dimension allows for reflection of the visual impression and deepening of the state of contemplation. For the resulting compositions, the passing scenery can be considered the score. "Re-transcription" of this score to an image gives a panoramic overview over the complete journey and exhibits some interesting effects caused by the movement of the train, such as compression and stretching of passing objects. In addition to intensifying the experience of a train journey, sound/tracks permits to persistently capture and archive the fleeting impressions of journey and composition and allows for re-experiencing the trip both visually and acoustically at a later point. Peter Knees, Tim Pohle, Gerhard Widmer |
ACM Multimedia | 3 |
| 2008 | Sound/tracks: real-time synaesthetic sonification and visualisation of passing landscapesabstractWhen travelling on a train, many people enjoy looking out of the window at the landscape passing by. We present sound/tracks, an application that translates the perceived movement of the scenery and other visual impressions, such as passing trains, into music. The continuously changing view outside the window is captured with a camera and translated into MIDI events that are replayed instantaneously. This allows for a reflection of the visual impression, adding a sound dimension to the visual experience and deepening the state of contemplation. The application is intended to be run on both mobile phones (with built-in camera) and on laptops (with a connected Web-cam). We propose and discuss different approaches to translating the video signal into an audio stream, present different application scenarios, and introduce a method to visualise the dynamics of complete train journeys by re-transcribing the captured video frames used to generate the music. Tim Pohle, Peter Knees, Gerhard Widmer |
ACM Multimedia | 3 |
| 2008 | Using string kernels to identify famous performers from their playing style
Craig Saunders, David R. Hardoon, John Shawe-Taylor, Gerhard Widmer |
Intell. Data Anal. | 4 |
| 2007 | Evaluating Low-Level Features for Beat Classification and TrackingabstractIn this paper, we address the question of which low-level acoustical features are the most suitable for identifying music beats computationally. We consider 172 features computed on consecutive signal frames and systematically evaluate their individual value in the task of providing reliable cues for the presence and localisation of beats in music signals. We compare two ways of evaluating features: their accuracy in a song-specific classification task (classifying beats vs nonbeats) and their performance as a front-end to a beat tracking system. Fabien Gouyon, Simon Dixon, Gerhard Widmer |
ICASSP (4) | 3 |
| 2007 | Towards a Computational Model of Melody Identification in Polyphonic Music
Søren Tjagvad Madsen, Gerhard Widmer |
IJCAI | 2 |
| 2007 | One-touch access to music on mobile devicesabstractWe present an approach that offers the user a convenient and meaningful way to access her music on a mobile device. By exploiting information on acoustic similarity and community-based music labels, a music collection is automatically structured and described to allow for easy orientation and navigation within the collection. To this end, the complete collection is arranged along a circular playlist path such that similar sounding pieces are grouped together. As a consequence, regions of musical styles emerge. Furthermore, we propose two approaches to derive informative descriptors that are displayed on the different regions, allowing an overview of the whole collection at a glance. For demonstration, we implemented our prototype interface on an Apple iPod. Dominik Schnitzer, Tim Pohle, Peter Knees, Gerhard Widmer |
MUM | 4 |
| 2007 | A music search engine built upon audio-based and web-based similarity measuresabstractAn approach is presented to automatically build a search engine for large-scale music collections that can be queried through natural language. While existing approaches depend on explicit manual annotations and meta-data assigned to the individual audio pieces, we automatically derive descriptions by making use of methods from Web Retrieval and Music Information Retrieval. Based on the ID3 tags of a collection of mp3 files, we retrieve relevant Web pages via Google queries and use the contents of these pages to characterize the music pieces and represent them by term vectors. By incorporating complementary information about acous tic similarity we are able to both reduce the dimensionality of the vector space and improve the performance of retrieval, i.e. the quality of the results. Furthermore, the usage of audio similarity allows us to also characterize audio pieces when there is no associated information found on the Web. Peter Knees, Tim Pohle, Markus Schedl, Gerhard Widmer |
SIGIR | 4 |
| 2007 | "Reinventing the Wheel": A Novel Approach to Music Player InterfacesabstractWe present a novel interface to (portable) music players that benefit from intelligently structured collections of audio files. For structuring, we calculate similarities between every pair of songs and model a travelling salesman problem (TSP) that is solved to obtain a playlist (i.e., the track ordering during playback) where the average distance between consecutive pieces of music is minimal according to the similarity measure. The similarities are determined using both audio signal analysis of the music tracks and Web-based artist profile comparison. Indeed, we show how to enhance the quality of the well-established methods based on audio signal processing with features derived from Web pages of music artists. Using TSP allows for creating circular playlists that can be easily browsed with a wheel as input device. We investigate the usefulness of four different TSP algorithms for this purpose. For evaluating the quality of the generated playlists, we apply a number of quality measures to two real-world music collections. It turns out that the proposed combination of audio and text-based similarity yields better results than the initial approach based on audio data only. We implemented an audio player as Java applet to demonstrate the benefits of our approach. Furthermore, we present the results of a small user study conducted to evaluate the quality of the generated playlists Tim Pohle, Peter Knees, Markus Schedl, Elias Pampalk, Gerhard Widmer |
IEEE Trans. Multim. | 5 |
| 2006 | Towards Automatic Retrieval of Album Covers
Markus Schedl, Peter Knees, Tim Pohle, Gerhard Widmer |
ECIR | 4 |
| 2006 | An innovative three-dimensional user interface for exploring music collections enrichedabstractWe present a novel, innovative user interface to music repositories. Given an arbitrary collection of digital music files, our system creates a virtual landscape which allows the user to freely navigate in this collection. This is accomplished by automatically extracting features from the audio signal and training a Self-Organizing Map (SOM) on them to form clusters of similar sounding pieces of music. Subsequently, a Smoothed Data Histogram (SDH) is calculated on the SOM and interpreted as a three-dimensional height profile. This height profile is visualized as a three-dimensional island landscape containing the pieces of music. While moving through the terrain, the closest sounds with respect to the listener's current position can be heard. This is realized by anisotropic auralization using a 5.1 surround sound model. Additionally, we incorporate knowledge extracted automatically from the web to enrich the landscape with semantic information. More precisely, we display words and related images that describe the heard music on the landscape to support the exploration. Peter Knees, Markus Schedl, Tim Pohle, Gerhard Widmer |
ACM Multimedia | 4 |
| 2006 | Relational IBL in classical music
Asmir Tobudic, Gerhard Widmer |
Mach. Learn. | 2 |
| 2006 | Guest editorial: Machine learning in and for music
Gerhard Widmer |
Mach. Learn. | 1 |
| 2006 | Guest Editorial: Machine learning in and for music
Gerhard Widmer |
Mach. Learn. | 1 |
| 2005 | Learning to Play Like the Great Pianists
Asmir Tobudic, Gerhard Widmer |
IJCAI | 2 |
| 2005 | Why Computers Need to Learn About Music
Gerhard Widmer |
ILP | 1 |
| 2005 | Interactive Poster: Using CoMIRVA for Visualizing Similarities Between Music Artists
Markus Schedl, Peter Knees, Gerhard Widmer |
IEEE Visualization | 3 |
| 2005 | Automatic identification of music performers with learning ensembles
Efstathios Stamatatos, Gerhard Widmer |
Artif. Intell. | 2 |
| 2004 | Automatic Recognition of Famous Artists by Machine
Gerhard Widmer, Patrick Zanon |
ECAI | 1 |
| 2004 | Using String Kernels to Identify Famous Performers from Their Playing Style
Craig Saunders, David R. Hardoon, John Shawe-Taylor, Gerhard Widmer |
ECML | 4 |
| 2004 | A new approach to hierarchical clustering and structuring of data with Self-Organizing Maps
Elias Pampalk, Gerhard Widmer, Alvin Chan Toong Shoon |
Intell. Data Anal. | 2 |
| 2003 | Playing Mozart Phrase by Phrase
Asmir Tobudic, Gerhard Widmer |
ICCBR | 2 |
| 2003 | Relational IBL in Music with a New Structural Similarity Measure
Asmir Tobudic, Gerhard Widmer |
ILP | 2 |
| 2003 | Visualizing changes in the structure of data for exploratory feature selectionabstractUsing visualization techniques to explore and understand high-dimensional data is an efficient way to combine human intelligence with the immense brute force computation power available nowadays. Several visualization techniques have been developed to study the cluster structure of data, i.e., the existence of distinctive groups in the data and how these clusters are related to each other. However, only few of these techniques lend themselves to studying how this structure changes if the features describing the data are changed. Understanding this relationship between the features and the cluster structure means understanding the features themselves and is thus a useful tool in the feature extraction phase.In this paper we present a novel approach to visualizing how modification of the features with respect to weighting or normalization changes the cluster structure. We demonstrate the application of our approach in two music related data mining projects. Elias Pampalk, Werner Goebl, Gerhard Widmer |
KDD | 3 |
| 2003 | Discovering simple rules in complex data: A meta-learning algorithm and some surprising musical discoveries
Gerhard Widmer |
Artif. Intell. | 1 |
| 2002 | In Search of the Horowitz Factor: Interim Report on a Musical Discovery Project
Gerhard Widmer |
ALT | 1 |
| 2002 | In Search of the Horowitz Factor: Interim Report on a Musical Discovery Project
Gerhard Widmer |
Discovery Science | 1 |
| 2002 | Music Performer Recognition Using an Ensemble of Simple Classifiers
Efstathios Stamatatos, Gerhard Widmer |
ECAI | 2 |
| 2002 | Towards a Simple Clustering Criterion Based on Minimum Length Encoding
Marcus-Christopher Ludl, Gerhard Widmer |
ECML | 2 |
| 2002 | Transformation-Based Regression
Björn Bringmann, Stefan Kramer 0001, Friedrich Neubarth, Hannes Pirker, Gerhard Widmer |
ICML | 5 |
| 2001 | Discovering Strong Principles of Expressive Music Performance with the PLCG Rule Learning Strategy
Gerhard Widmer |
ECML | 1 |
| 2001 | The Musical Expression Project: A Challenge for Machine Learning and Knowledge Discovery
Gerhard Widmer |
ECML | 1 |
| 2001 | The Musical Expression Project: A Challenge for Machine Learning and Knowledge Discovery
Gerhard Widmer |
PKDD | 1 |
| 2001 | Prediction of Ordinal Classes Using Regression Trees
Stefan Kramer 0001, Gerhard Widmer, Bernhard Pfahringer, Michael de Groeve |
Fundam. Informaticae | 2 |
| 2000 | Relative Unsupervised Discretization for Regresseion Problems
Marcus-Christopher Ludl, Gerhard Widmer |
ECML | 2 |
| 2000 | Prediction of Ordinal Classes Using Regression Trees
Stefan Kramer 0001, Gerhard Widmer, Bernhard Pfahringer, Michael de Groeve |
ISMIS | 2 |
| 2000 | Relative Unsupervised Discretization for Association Rule Mining
Marcus-Christopher Ludl, Gerhard Widmer |
PKDD | 2 |
| 2000 | Searching for Patterns in Political Event Sequences: Experiments with the Keds DatabaseabstractThis paper presents an empirical study on the possibility of discovering interesting event sequences and sequential rules in a large database of international political events. A data mining algorithm first presented by Mannila and Toivonen (1996), has been implemented and extended, which is able to search for generalized episodes in such event databases. Experiments conducted with this algorithm on the Kansas Event Data System (KEDS) database, an event data set covering interactions between countries in the Persian Gulf region, are described. Some qualitative and quantitative results are reported, and experiences with strategies for reducing the problem complexity and focusing on the search on interesting subsets of events are described. Klaus Kovar, Johannes Fürnkranz, Johann Petrak, Bernhard Pfahringer, Robert Trappl, Gerhard Widmer |
Cybern. Syst. | 6 |
| 1998 | Guest Editors' Introduction
Gerhard Widmer, Miroslav Kubat |
Mach. Learn. | 1 |
| 1997 | Tracking Context Changes through Meta-Learning
Gerhard Widmer |
Mach. Learn. | 1 |
| 1996 | What Is It That Makes It a Horowitz? Empirical Musicology via Machine Learning
Gerhard Widmer |
ECAI | 1 |
| 1996 | Recognition and Exploitation of Contextual CLues via Incremental Meta-Learning
Gerhard Widmer |
ICML | 1 |
| 1996 | Learning in the Presence of Concept Drift and Hidden Contexts
Gerhard Widmer, Miroslav Kubat |
Mach. Learn. | 1 |
| 1995 | Adapting to Drift in Continuous Domains (Extended Abstract)
Miroslav Kubat, Gerhard Widmer |
ECML | 2 |
| 1994 | The Synergy of Music Theory and Al: Learning Multi-Level Expressive Interpretation
Gerhard Widmer |
AAAI | 1 |
| 1994 | Combining Robustness and Flexibility in Learning Drifting Concepts
Gerhard Widmer |
ECAI | 1 |
| 1994 | Incremental Reduced Error Pruning
Johannes Fürnkranz, Gerhard Widmer |
ICML | 2 |
| 1993 | Effective Learning in Dynamic Environments by Explicit Context Tracking
Gerhard Widmer, Miroslav Kubat |
ECML | 1 |
| 1993 | Automatic knowledge base refinement: learning from examples and deep knowledge in rheumatology
Gerhard Widmer, Werner Horn, Bernhard Nagele |
Artif. Intell. Medicine | 1 |
| 1992 | Learning Flexible Concepts from Streams of Examples: FLORA 2
Gerhard Widmer, Miroslav Kubat |
ECAI | 1 |
| 1989 | A Tight Integration of Deductive Learning
Gerhard Widmer |
ML | 1 |