VLDB 2026 Research / reviewers in the wild / expert
Bob L. T. Sturm
dblp:85/3478
· DBLP profile ↗
27ranked-venue papers
11as first author
7since 2021 · last 2026
0000-0003-2549-6367ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 15 · 7 first-author · 3 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Charting Creative Journeys through Musical Performance with AIabstractWe chart the creative journeys of two groups of musicians as they explored how to engage with an AI system developed for the live performance of Irish Traditional Dance Music. A group of expert Irish musicians participated in an extended workshop to refine the system while a professional contemporary folk duo composed and rehearsed material for a public show. Drawing on the standard definition of creativity, we chart their creative journeys across a landscape we define by the two dimensions of originality and effectiveness. We reveal the importance of authenticity to a traditional practice as an opposing force to creativity and consider how the tension between these shaped their musical journeys. We generalise five strategies that move creative practitioners across this landscape in different directions: proving your chops, finding your place, perfecting the nuances, innovating, and glitching. Marco Amerotti, Steve Benford, Adrian Hazzard, Juan Pablo Martinez-Avila, Paul Tennent, Bob L. T. Sturm |
Creativity & Cognition | 6 |
| 2025 | (Mis)Communicating with our AI SystemsabstractExplainable Artificial Intelligence (XAI) is a discipline concerned with understanding predictions of AI systems. What is ultimately desired from XAI methods is for an AI system to link its input and output in a way that is interpretable with reference to the environment in which it is applied. A variety of methods have been proposed, but we argue in this paper that what has yet to be considered is miscommunication: the failure to convey and/or interpret an explanation accurately. XAI can be seen as a communication process and thus looking at how humans explain things to each other can provide guidance to its application and evaluation. We motivate a specific model of communication to help identify essential components of the process, and show the critical importance for establishing common ground, i.e., shared mutual knowledge, beliefs, and assumptions of the participants communicating. Laura Cros Vila, Bob L. T. Sturm |
CHI | 2 |
| 2025 | Exploring the Expressive Space of an Articulatory Vocal Modal using Quality-Diversity Optimization with Multimodal EmbeddingsabstractKnowing which sounds can be produced by a simulated vocal model and how they are connected to its articulatory behavior is not trivial. Being able to map this out can be interesting for applications that make use of the extended capabilities of a voice, e.g., singing or vocal imitations. We present a method that achieves this for a state-of-the-art articulatory vocal model (VocalTractLab) by combining it with a recent Quality-Diversity algorithm (CMA-MAE) and audio embeddings obtained through a multi-modal pretrained model (CLAP). The text-capabilities of CLAP make it possible to steer the exploration through a text prompt. We show that the method explores more efficiently than a random sampling baseline, covering more of the measure space and achieving higher objective scores. We provide several listening examples and the source code for a scalable implementation. Joris Grouwels, Nicolas Jonason, Bob L. T. Sturm |
GECCO | 3 |
| 2024 | Investigating the Viability of Masked Language Modeling for Symbolic Music Generation in abc-notation
Luca Casini, Nicolas Jonason, Bob L. T. Sturm |
EvoMUSART | 3 |
| 2024 | The Chordinator: Modeling Music Harmony by Implementing Transformer Networks and Token Strategies
David Dalmazzo, Ken Déguernel, Bob L. T. Sturm |
EvoMUSART | 3 |
| 2023 | Bias in Favour or Against Computational Creativity: A Survey and Reflection on the Importance of Socio-cultural Context in its Evaluation
Ken Déguernel, Bob L. T. Sturm |
ICCC | 2 |
| 2022 | Tradformer: A Transformer Model of Traditional Music TranscriptionsabstractWe explore the transformer neural network architecture for modeling music, specifically Irish and Swedish traditional dance music. Given the repetitive structures of these kinds of music, the transformer should be as successful with fewer parameters and complexity as the hitherto most successful model, a vanilla long short-term memory network. We find that achieving good performance with the transformer is not straightforward, and careful consideration is needed for the sampling strategy, evaluating intermediate outputs in relation to engineering choices, and finally analyzing what the model learns. We discuss these points with several illustrations, providing reusable insights for engineering other music generation systems. We also report the high performance of our final transformer model in a competition of music generation systems focused on a type of Swedish dance. Luca Casini, Bob L. T. Sturm |
IJCAI | 2 |
| 2020 | Reliable Local Explanations for Machine ListeningabstractOne way to analyse the behaviour of machine learning models is through local explanations that highlight input features that maximally influence model predictions. Sensitivity analysis, which involves analysing the effect of input perturbations on model predictions, is one of the methods to generate local explanations. Meaningful input perturbations are essential for generating reliable explanations, but there exists limited work on what such perturbations are and how to perform them. This work investigates these questions in the context of machine listening models that analyse audio. Specifically, we use a state-of-the-art deep singing voice detection (SVD) model to analyse whether explanations from SoundLIME (a local explanation method) are sensitive to how the method perturbs model inputs. The results demonstrate that SoundLIME explanations are sensitive to the content in the occluded input regions. We further propose and demonstrate a novel method for quantitatively identifying suitable content type(s) for reliably occluding inputs of machine listening models. The results for the SVD model suggest that the average magnitude of input mel-spectrogram bins is the most suitable content type for temporal explanations. Saumitra Mishra, Emmanouil Benetos, Bob L. T. Sturm, Simon Dixon |
IJCNN | 3 |
| 2020 | Dataset Artefacts in Anti-Spoofing Systems: A Case Study on the ASVspoof 2017 BenchmarkabstractThe Automatic Speaker Verification Spoofing and Countermeasures Challenges motivate research in protecting speech biometric systems against a variety of different access attacks. The 2017 edition focused on replay spoofing attacks, and involved participants building and training systems on a provided dataset (ASVspoof 2017). More than 60 research papers have so far been published with this dataset, but none have sought to answer why countermeasures appear successful in detecting spoofing attacks. This article shows how artefacts inherent to the dataset may be contributing to the apparent success of published systems. We first inspect the ASVspoof 2017 dataset and summarize various artefacts present in the dataset. Second, we demonstrate how countermeasure models can exploit these artefacts to appear successful in this dataset. Third, for reliable and robust performance estimates on this dataset we propose discarding nonspeech segments and silence before and after the speech utterance during training and inference. We create speech start and endpoint annotations in the dataset and demonstrate how using them helps countermeasure models become less vulnerable from being manipulated using artefacts found in the dataset. Finally, we provide several new benchmark results for both frame-level and utterance-level models that can serve as new baselines on this dataset. Bhusan Chettri, Emmanouil Benetos, Bob L. T. Sturm |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Ensemble Models for Spoofing Detection in Automatic Speaker VerificationabstractDetecting spoofing attempts of automatic speaker verification (ASV) systems is challenging, especially when using only one modeling approach. For robustness, we use both deep neural networks and traditional machine learning models and combine them as ensemble models through logistic regression. They are trained to detect logical access (LA) and physical access (PA) attacks on the dataset released as part of the ASV Spoofing and Countermeasures Challenge 2019. We propose dataset partitions that ensure different attack types are present during training and validation to improve system robustness. Our ensemble model outperforms all our single models and the baselines from the challenge for both attack types. We investigate why some models on the PA dataset strongly outperform others and find that spoofed recordings in the dataset tend to have longer silences at the end than genuine ones. By removing them, the PA task becomes much more challenging, with the tandem detection cost function (t-DCF) of our best single model rising from 0.1672 to 0.5018 and equal error rate (EER) increasing from 5.98% to 19.8% on the development set. Bhusan Chettri, Daniel Stoller, Veronica Morfi, Marco A. Martínez Ramírez, Emmanouil Benetos, Bob L. T. Sturm |
INTERSPEECH | 6 |
| 2018 | A Deeper Look at Gaussian Mixture Model Based Anti-Spoofing SystemsabstractA “replay attack” involves replaying pre-recorded speech of an enrolled speaker to bypass an automatic speaker verification system. The 2017 ASVspoof Challenge focused on this kind of attack. In this paper, we describe our evaluation work after this challenge. First, we study the effectiveness of Gaussian Mixture Model (GMM) systems using six different hand-crafted features for detecting a replay attack. Second, we take a deeper look at these GMM systems and perform a frame-level analysis of log likelihoods. Our analysis shows how system performance can depend on a simple class-dependent cue in the dataset: initial silence frames of zeros appear in the genuine signals but missing in the spoofed version. Third, we show how we can fool these systems using this cue. For example, we find the equal error rate (EER) of one GMM system dramatically rises from 14.82 to 44.44 when we add the cue to the evaluation data. Finally, we explore whether this problem can be mitigated by pre-processing the 2017 ASV spoof Challenge dataset. Bhusan Chettri, Bob L. T. Sturm |
ICASSP | 2 |
| 2018 | Analysing The Predictions Of a CNN-Based Replay Spoofing Detection SystemabstractPlaying recorded speech samples of an enrolled speaker - “replay attack” - is a simple approach to bypass an automatic speaker verification (ASV) system. The vulnerability of ASV systems to such attacks has been acknowledged and studied, but there has been no research into what spoofing detection systems are actually learning to discriminate. In this paper, we analyse the local behaviour of a replay spoofing detection system based on convolutional neural networks (CNNs) adapted from a state-of-the-art CNN (LCNNFFT) submitted at the ASVspoof 2017 challenge. We generate temporal and spectral explanations for predictions of the model using the SLIME algorithm. Our findings suggest that in most instances of spoofing the model is using information in the first 400 milliseconds of each audio instance to make the class prediction. Knowledge of the characteristics that spoofing detection systems are exploiting can help build less vulnerable ASV systems, other spoofing detection systems, as well as better evaluation databases1. Bhusan Chettri, Saumitra Mishra, Bob L. T. Sturm, Emmanouil Benetos |
SLT | 3 |
| 2016 | The "Beyond the Fence" Musical and "Computer Says Show" Documentary
Simon Colton, Maria Teresa Llano, Rose Hepworth, John William Charnley, Catherine V. Gale, Archie Baron, François Pachet, Pierre Roy, Pablo Gervás, Nick Collins, Bob L. T. Sturm, Tillman Weyde, Daniel Wolff, James Robert Lloyd |
ICCC | 11 |
| 2015 | Deep Learning and Music AdversariesabstractAn adversary is an agent designed to make a classification system perform in some particular way, e.g., increase the probability of a false negative. Recent work builds adversaries for deep learning systems applied to image object recognition, exploiting the parameters of the system to find the minimal perturbation of the input image such that the system misclassifies it with high confidence. We adapt this approach to construct and deploy an adversary of deep learning systems applied to music content analysis. In our case, however, the system inputs are magnitude spectral frames, which require special care in order to produce valid input audio signals from network- derived perturbations . For two different train-test partitionings of two benchmark datasets, and two different architectures , we find that this adversary is very effective. We find that convolutional architectures are more robust compared to systems based on a majority vote over individually classified audio frames. Furthermore , we experiment with a new system that integrates an adversary into the training loop, but do not find that this improves the resilience of the system to new adversaries. Corey Kereliuk, Bob L. T. Sturm, Jan Larsen |
IEEE Trans. Multim. | 2 |
| 2014 | A Simple Method to Determine if a Music Information Retrieval System is a "Horse"abstractWe propose and demonstrate a simple method to explain the figure of merit (FoM) of a music information retrieval (MIR) system evaluated in a dataset, specifically, whether the FoM comes from the system using characteristics confounded with the “ground truth” of the dataset. Akin to the controlled experiments designed to test the supposed mathematical ability of the famous horse “Clever Hans,” we perform two experiments to show how three state-of-the-art MIR systems produce excellent FoM in spite of not using musical knowledge. This provides avenues for improving MIR systems, as well as their evaluation. We make available a reproducible research package so that others can apply the same method to evaluating other MIR systems. Bob L. T. Sturm |
IEEE Trans. Multim. | 1 |
| 2013 | Behavior of greedy sparse representation algorithms on nested supportsabstractIn this work, we study the links between the recovery properties of sparse signals for Orthogonal Matching Pursuit (OMP) and the whole General MP class over nested supports. We show that the optimality of those algorithms is not locally nested: there is a dictionary and supports I and J with J included in I such that OMP will recover all signals of support I, but not all signals of support J. We also show that the optimality of OMP is globally nested: if OMP can recover all s-sparse signals, then it can recover all s'-sparse signals with s' smaller than s. We also provide a tighter version of Donoho and Elad's spark theorem, which allows us to complete Tropp's proof that sparse representation algorithms can only be optimal for all s-sparse signals if s is strictly lower than half the spark of the dictionary. Boris Mailhé, Bob L. T. Sturm, Mark D. Plumbley |
ICASSP | 2 |
| 2013 | On music genre classification via compressive samplingabstractRecent work [1] combines low-level acoustic features and random projection (referred to as “compressed sensing” in [1]) to create a music genre classification system showing an accuracy among the highest reported for a benchmark dataset. This not only contradicts previous findings that suggest low-level features are inadequate for addressing high-level musical problems, but also that a random projection of features can improve classification. We reproduce this work and resolve these contradictions. Bob L. T. Sturm |
ICME | 1 |
| 2013 | Music genre recognition with risk and rejectionabstractWe explore risk and rejection for music genre recognition (MGR) within the minimum risk framework of Bayesian classification. In this way, we attempt to give an MGR system knowledge that some misclassifications are worse than others, and that deferring classification to an expert may be a better option than forcing a label under high uncertainty. Our experiments show this approach to have some success with respect to reducing false positives and negatives. Bob L. T. Sturm |
ICME | 1 |
| 2013 | Classification accuracy is not enough - On the evaluation of music genre recognition systemsabstractWe argue that an evaluation of system behavior at the level of the music is required to usefully address the fundamental problems of music genre recognition (MGR), and indeed other tasks of music information retrieval, such as autotagging. A recent review of works in MGR since 1995 shows that most (82 %) measure the capacity of a system to recognize genre by its classification accuracy. After reviewing evaluation in MGR, we show that neither classification accuracy, nor recall and precision, nor confusion tables, necessarily reflect the capacity of a system to recognize genre in musical signals. Hence, such figures of merit cannot be used to reliably rank, promote or discount the genre recognition performance of MGR systems if genre recognition (rather than identification by irrelevant confounding factors) is the objective. This motivates the development of a richer experimental toolbox for evaluating any system designed to intelligently extract information from music signals. Bob L. T. Sturm |
J. Intell. Inf. Syst. | 1 |
| 2013 | Revisiting Inter-Genre SimilarityabstractWe revisit the idea of “inter-genre similarity” (IGS) for machine learning in general, and music genre recognition in particular. We show analytically that the probability of error for IGS is higher than naive Bayes classification with zero-one loss (NB). We show empirically that IGS does not perform well, even for data that satisfies all its assumptions. Bob L. T. Sturm, Fabien Gouyon |
IEEE Signal Process. Lett. | 1 |
| 2013 | On Theorem 10 in "On Polar Polytopes and the Recovery of Sparse Representations" [Sep 07 3188-3195]abstractIt is shown that Theorem 10 (Non-Nestedness of ERC) in [Plumbley, IEEE Trans. Inf. Theory, vol. 53, pp. 3188-3195, Sep. 2007] neglects the derivations of the exact recovery conditions (ERCs) of constrained l1-minimization (BP) and orthogonal matching pursuit. This means that it does not reflect the recovery properties of these algorithms. Furthermore, an ERC of BP more general than that in [Tropp, IEEE Trans. Inf. Theory, vol. 50, pp. 2231-2242, Oct. 2004] is shown. Bob L. T. Sturm, Boris Mailhé, Mark D. Plumbley |
IEEE Trans. Inf. Theory | 1 |
| 2011 | Recursive nearest neighbor search in a sparse and multiscale domain for comparing audio signals
Bob L. T. Sturm, Laurent Daudet |
Signal Process. | 1 |
| 2010 | Incorporating scale information with cepstral features: Experiments on musical instrument recognition
Marcela Morvidone, Bob L. T. Sturm, Laurent Daudet |
Pattern Recognit. Lett. | 2 |
| 2010 | Sparse Approximation and the Pursuit of Meaningful Signal Models With Interference AdaptationabstractIn the pursuit of a sparse signal model, mismatches between the signal and the dictionary, as well as atoms poorly selected by the decomposition process, can diminish the efficiency and meaningfulness of the resulting representation. These problems increase the number of atoms needed to model a signal for a given error, and they obscure the relationships between signal content and the elements of the model. To increase the efficiency and meaningfulness of a signal model built by an iterative descent pursuit, such as matching pursuit (MP), we propose integrating into its atom selection criterion a measure ofinterferencebetween an atom and the model. We define interference and illustrate how it describes the contribution of an atom to modeling a signal. We show that for any nontrivial signal, the convergent model created by MP must have as much destructive as constructive interference, i.e., MP cannot avoid correction in the signal model. This is not necessarily a shortcoming of orthogonal variants of MP, such as orthogonal MP (OMP). We derive interference-adaptive iterative descent pursuits and show how these can build signal models that better fit the signal locally, and reduce the corrections made in a signal model. Compared with MP and its orthogonal variants, our experimental results not only show an increase in model efficiency, but also a clearer correspondence between the signal and the atoms of a representation. Bob L. T. Sturm, John J. Shynk |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Agglomerative clustering in sparse atomic decompositions of audio signalsabstractWe present a correlation-based algorithm for the agglomerative clustering of atoms in sparse atomic decompositions of audio signals. Our goal is to demonstrate useful relationships between elements of the decomposition and the content of the original signal, for such purposes as analysis and modification. We evaluate the performance of the agglomeration algorithm using decompositions of synthetic and real audio signals, and discuss possible extensions of this work. Bob L. T. Sturm, John J. Shynk, Steffen Gauglitz |
ICASSP | 1 |
| 2008 | Dark Energy in Sparse Atomic EstimationsabstractSparse overcomplete methods, such as matching pursuit, attempt to find an efficient estimation of a signal using terms (atoms) selected from an overcomplete dictionary. In some cases, atoms can be selected that have energy in regions of the signal that have no energy. Other atoms are then used to destructively interfere with these terms in order to preserve the original waveform. Because some terms may even ldquodisappearrdquo in the reconstruction, we refer to the destructive and constructive interference between the atoms of a sparse atomic estimation as ldquodark energy.rdquo In this paper, we formally define dark energy for matching pursuit, explore its properties, and present empirical results for decompositions of audio signals. This paper demonstrates that dark energy is a useful measure of the interference between the terms of a sparse atomic estimation and might provide information for the decomposition process. Bob L. T. Sturm, John J. Shynk, Laurent Daudet, Curtis Roads |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Matching Pursuit Decompositions of Non-Noisy Speech Signals Using Several DictionariesabstractMatching pursuit (MP) provides a way to expand signals in terms of any set of time-limited functions, or atoms, called a dictionary. These decompositions are finding use in signal analysis and coding. It has been shown that a dictionary should be designed carefully, but its effects on decomposition have not been studied in detail. We look at the effects of dictionaries on the decomposition of non-noisy speech signals using MP, by five dictionaries. It is found that Gabor atoms work sufficiently well, and have fewer adverse effects in reconstruction compared to the other dictionaries. For a reconstruction to sound perceptually close to the original, a rate of 3000 atoms per second (aps) on average is required. At rates as low as 400 aps the speech remains intelligible. Finally, the use of decompositions to visualize time-frequency distributions of speech is explored. Bob L. T. Sturm, Jerry D. Gibson |
ICASSP (3) | 1 |