EDBT 2026 Demo / reviewers in the wild / expert
Simon Dixon
dblp:11/7013 · also Simon E. Dixon
· DBLP profile ↗
60ranked-venue papers
5as first author
15since 2021 · last 2026
0000-0002-6098-481XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 21 · 5 first-author · 2 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reconstruction Meets Prediction: Complementary SSL Objectives for Music Transcription
Mary Pilataki, Matthias Mauch, Simon Dixon |
IEEE Signal Process. Lett. | 3 |
| 2025 | LLaQo: Towards a Query-Based Coach in Expressive Music Performance AssessmentabstractResearch in music understanding has extensively explored composition-level attributes such as key, genre, and instrumentation through advanced representations, leading to cross-modal applications using large language models. However, aspects of musical performance such as stylistic expression and technique remain underexplored, along with the potential of using large language models to enhance educational outcomes with customized feedback. To bridge this gap, we introduce LLaQo, a Large Language Query-based music coach that leverages audio language modeling to provide detailed and formative assessments of music performances. We also introduce instruction-tuned query-response datasets that cover a variety of performance dimensions from pitch accuracy to articulation, as well as contextual performance understanding (such as difficulty and performance techniques). Utilizing AudioMAE encoder and Vicuna-7b LLM backend, our model achieved state-of-the-art (SOTA) results in predicting teachers’ performance ratings, as well as in identifying piece difficulty and playing techniques. Textual responses from LLaQo was moreover rated significantly higher compared to other baseline models in a user study using audio-text matching. Our proposed model can thus provide informative answers to open-ended questions related to musical performance from audio data. Huan Zhang 0001, Vincent K. M. Cheung, Hayato Nishioka, Simon Dixon, Shinichi Furuya |
ICASSP | 4 |
| 2024 | Posterior Variance-Parameterised Gaussian Dropout: Improving Disentangled Sequential Autoencoders for Zero-Shot Voice ConversionabstractThe class of disentangled sequential auto-encoders factorises speech into time-invariant (global) and time-variant (local) representations for speaker identity and linguistic content, respectively. Many of the existing models employ this assumption to tackle zero-shot voice conversion (VC), which converts speaker characteristics of any given utterance to any novel speakers while preserving the linguistic content. However, balancing capacity between the two representations is intricate, as the global representation tends to collapse due to its lower information capacity along the time axis than that of the local representation. We propose a simple and effective dropout technique that applies an information bottleneck to the local representation via multiplicative Gaussian noise, in order to encourage the usage of the global one. We endow existing zero-shot VC models with the proposed method and show significant improvements in speaker conversion in terms of speaker verification acceptance rate and comparable or better intelligibility measured in character error rate. Yin-Jyun Luo, Simon Dixon |
ICASSP | 2 |
| 2024 | Unsupervised Pitch-Timbre Disentanglement of Musical Instruments Using a Jacobian Disentangled Sequential AutoencoderabstractDisentangled representation learning seeks to align individual dimensions or separate groups of coordinates of latent factors with attributes of observed data such that perturbing certain latent factors uniquely changes particular attributes. A main challenge in unsupervised disentanglement using autoencoders is that strong regularisation, while necessary for consistent disentanglement, comes at the expense of accurate data reconstruction. To address this, we introduce a teacher-student framework that incorporates a variational sequential autoencoder and a Jacobian constraint that regularises the variation of observations relative to latent factors. In real-world audio recordings of musical instruments, our approach outperforms a state-of-the-art method in both sampling quality and unsupervised pitch-timbre disentanglement. Yin-Jyun Luo, Sebastian Ewert, Simon Dixon |
ICASSP | 3 |
| 2024 | High Resolution Guitar Transcription Via Domain AdaptationabstractAutomatic music transcription (AMT) has achieved high accuracy for piano due to the availability of large, high-quality datasets such as MAESTRO and MAPS, but comparable datasets are not yet available for other instruments. In recent work, however, it has been demonstrated that aligning scores to transcription model activations can produce high quality AMT training data for instruments other than piano. Focusing on the guitar, we refine this approach to training on score data using a dataset of commercially available score-audio pairs. We propose the use of a high-resolution piano transcription model to train a new guitar transcription model. The resulting model obtains state-of-the-art transcription results on GuitarSet in a zero-shot context, improving on previously published methods. Xavier Riley, Drew Edwards, Simon Dixon |
ICASSP | 3 |
| 2024 | MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models
Yixiao Zhang 0002, Yukara Ikemiya, Gus Xia, Naoki Murata, Marco A. Martínez Ramírez, Wei-Hsiang Liao 0001, Yuki Mitsufuji, Simon Dixon |
IJCAI | 8 |
| 2024 | Pitch-aware generative pretraining improves multi-pitch estimation with scarce data
Mary Pilataki, Matthias Mauch, Simon Dixon |
MMAsia | 3 |
| 2024 | A Data-Driven Analysis of Robust Automatic Piano TranscriptionabstractAlgorithms for automatic piano transcription have improved dramatically in recent years due to new datasets and modeling techniques. Recent developments have focused primarily on adapting new neural network architectures, such as the Transformer and Perceiver, in order to yield more accurate systems. In this work, we study transcription systems from the perspective of their training data. By measuring their performance on out-of-distribution annotated piano data, we show how these models can severely overfit to acoustic properties of the training data. We create a new set of audio for the MAESTRO dataset, captured automatically in a professional studio recording environment via Yamaha Disklavier playback. Using various data augmentation techniques when training with the original and re-performed versions of the MAESTRO dataset, we achieve state-of-the-art note-onset accuracy of 88.4 F1-score on the MAPS dataset, without seeing any of its training data. We subsequently analyze these data augmentation techniques in a series of ablation studies to better understand their influence on the resulting models. Drew Edwards, Simon Dixon, Emmanouil Benetos, Akira Maezawa, Yuta Kusaka |
IEEE Signal Process. Lett. | 2 |
| 2023 | Disentangling the Horowitz Factor: Learning Content and Style From Expressive Piano PerformanceabstractIn the Western art music tradition, expressive piano performance consists of two kinds of information: the score, with pitch and timing expressed in simple musical units along with occasional expression instructions, and the performer’s interpretation of the score, involving variations in tempo, dynamics and articulation. In this paper, we present a novel framework for learning representations that disentangle musical content and performance style from expressive piano performances in an unsupervised manner. Our method is based on an extension of the vector-quantized variational autoencoder (VQ-VAE) with individual content and style branches, along with mutual information (MI) minimization techniques and self-supervising strategies. We performed experiments and ablation studies on the ATEPP dataset, a large set of automatically transcribed virtuosic piano performances with rich stylistic variations, and evaluated the content reconstruction and style discrimination in a style-transfer manner. Our experiments demonstrate that the model learnt separate latent variables that encode musical content (such as pitch and relative timing) and stylistic attributes, as generated samples align well with the content input with low note error rates (NER), and the 40-way style discrimination proxy task outperformed the baseline with top1 accuracy of 0.168. Huan Zhang 0001, Simon Dixon |
ICASSP | 2 |
| 2022 | Towards Robust Unsupervised Disentanglement of Sequential Data - A Case Study Using Music AudioabstractDisentangled sequential autoencoders (DSAEs) represent a class of probabilistic graphical models that describes an observed sequence with dynamic latent variables and a static latent variable. The former encode information at a frame rate identical to the observation, while the latter globally governs the entire sequence. This introduces an inductive bias and facilitates unsupervised disentanglement of the underlying local and global factors. In this paper, we show that the vanilla DSAE suffers from being sensitive to the choice of model architecture and capacity of the dynamic latent variables, and is prone to collapse the static latent variable. As a countermeasure, we propose TS-DSAE, a two-stage training framework that first learns sequence-level prior distributions, which are subsequently employed to regularise the model and facilitate auxiliary objectives to promote disentanglement. The proposed framework is fully unsupervised and robust against the global factor collapse problem across a wide range of model configurations. It also avoids typical solutions such as adversarial training which usually involves laborious parameter tuning, and domain-specific data augmentation. We conduct quantitative and qualitative evaluations to demonstrate its robustness in terms of disentanglement on both artificial and real-world music audio datasets. Yin-Jyun Luo, Sebastian Ewert, Simon Dixon |
IJCAI | 3 |
| 2022 | A Convolutional-Attentional Neural Framework for Structure-Aware Performance-Score SynchronizationabstractPerformance-score synchronization is an integral task in signal processing, which entails generating an accurate mapping between an audio recording of a performance and the corresponding musical score. Traditional synchronization methods compute alignment using knowledge-driven and stochastic approaches, and are typically unable to generalize well to different domains and modalities. We present a novel data-driven method for structure-aware performance-score synchronization. We propose a convolutional-attentional architecture trained with a custom loss based on time-series divergence. We conduct experiments for the audio-to-MIDI and audio-to-image alignment tasks pertained to different score modalities. We validate the effectiveness of our method via ablation studies and comparisons with state-of-the-art alignment approaches. We demonstrate that our approach outperforms previous synchronization methods for a variety of test settings across score modalities and acoustic conditions. Our method is also robust to structural differences between the performance and score sequences, which is a common limitation of standard alignment approaches. Ruchit Agrawal, Daniel Wolff, Simon Dixon |
IEEE Signal Process. Lett. | 3 |
| 2022 | The Jazz Ontology: A semantic model and large-scale RDF repositories for jazzabstractJazz is a musical tradition that is just over 100 years old; unlike in other Western musical traditions, improvisation plays a central role in jazz. Modelling the domain of jazz poses some ontological challenges due to specificities in musical content and performance practice, such as band lineup fluidity and importance of short melodic patterns for improvisation. This paper presents the Jazz Ontology – a semantic model that addresses these challenges. Additionally, the model also describes workflows for annotating recordings with melody transcriptions and for pattern search. The Jazz Ontology incorporates existing standards and ontologies such as FRBR and the Music Ontology. The ontology has been assessed by examining how well it supports describing and merging existing datasets and whether it facilitates novel discoveries in a music browsing application. The utility of the ontology is also demonstrated in a novel framework for managing jazz related music information. This involves the population of the Jazz Ontology with the metadata from large scale audio and bibliographic corpora (the Jazz Encyclopedia and the Jazz Discography). The resulting RDF datasets were merged and linked to existing Linked Open Data resources. These datasets are publicly available and are driving an online application that is being used by jazz researchers and music lovers for the systematic study of jazz. Polina Proutskova, Daniel Wolff, György Fazekas, Klaus Frieler, Frank Höger, Olga Velichkina, Gabriel Solis, Tillman Weyde, Martin Pfleiderer, Hélène C. Crayencour, Geoffroy Peeters, Simon Dixon |
J. Web Semant. | 12 |
| 2021 | Structure-Aware Audio-to-Score Alignment Using Progressively Dilated Convolutional Neural NetworksabstractThe identification of structural differences between a music performance and the score is a challenging yet integral step of audio-to-score alignment, an important subtask of music information retrieval. We present a novel method to detect such differences between the score and performance for a given piece of music using progressively dilated convolutional neural networks. Our method incorporates varying dilation rates at different layers to capture both short-term and long-term context, and can be employed successfully in the presence of limited annotated data. We conduct experiments on audio recordings of real performances that differ structurally from the score, and our results demonstrate that our models outperform standard methods for structure-aware audio-to-score alignment. Ruchit Agrawal, Daniel Wolff, Simon Dixon |
ICASSP | 3 |
| 2021 | Low Resource Audio-To-Lyrics Alignment from Polyphonic Music RecordingsabstractLyrics alignment in long music recordings can be memory exhaustive when performed in a single pass. In this study, we present a novel method that performs audio-to-lyrics alignment with a low memory consumption footprint regardless of the duration of the music recording. The proposed system first spots the anchoring words within the audio signal. With respect to these anchors, the recording is then segmented and a second-pass alignment is performed to obtain the word timings. We show that our audio-to-lyrics alignment system performs competitively with the state-of-the-art, while requiring much less computational resources. In addition, we utilize our lyrics alignment system to segment the music recordings into sentence-level chunks. Notably on the segmented recordings, we report the lyrics transcription scores on a number of benchmark test sets. Finally, our experiments highlight the importance of the source separation step for good performance on the transcription and alignment tasks. For reproducibility, we publicly share our code with the research community. Emir Demirel, Sven Ahlbäck, Simon Dixon |
ICASSP | 3 |
| 2021 | Adversarial Unsupervised Domain Adaptation for Harmonic-Percussive Source SeparationabstractThis letter addresses the problem of domain adaptation for the task of music source separation. Using datasets from two different domains, we compare the performance of a deep learning-based harmonic-percussive source separation model under different training scenarios, including supervised joint training using data from both domains and pre-training in one domain with fine-tuning in another. We propose an adversarial unsupervised domain adaptation approach suitable for the case where no labelled data (ground-truth source signals) from a target domain is available. By leveraging unlabelled data (only mixtures) from this domain, experiments show that our framework can improve separation performance on the new domain without losing any considerable performance on the original domain. The letter also introduces the Tap & Fiddle dataset, a dataset containing recordings of Scandinavian fiddle tunes along with isolated tracks for “foot-tapping” and “violin”. Carlos Lordelo, Emmanouil Benetos, Simon Dixon, Sven Ahlbäck, Patrik Ohlsson |
IEEE Signal Process. Lett. | 3 |
| 2020 | Training Generative Adversarial Networks from Incomplete Observations using Factorised Discriminators
Daniel Stoller, Sebastian Ewert, Simon Dixon |
ICLR | 3 |
| 2020 | Seq-U-Net: A One-Dimensional Causal U-Net for Efficient Sequence ModellingabstractConvolutional neural networks (CNNs) with dilated filters such as the Wavenet or the Temporal Convolutional Network (TCN) have shown good results in a variety of sequence modelling tasks. While their receptive field grows exponentially with the number of layers, computing the convolutions over very long sequences of features in each layer is time and memory-intensive, and prohibits the use of longer receptive fields in practice. To increase efficiency, we make use of the "slow feature" hypothesis stating that many features of interest are slowly varying over time. For this, we use a U-Net architecture that computes features at multiple time-scales and adapt it to our auto-regressive scenario by making convolutions causal. We apply our model ("Seq-U-Net") to a variety of tasks including language and audio generation. In comparison to TCN and Wavenet, our network consistently saves memory and computation time, with speed-ups for training and inference of over 4x in the audio generation experiment in particular, while achieving a comparable performance on real-world tasks. Daniel Stoller, Mi Tian 0001, Sebastian Ewert, Simon Dixon |
IJCAI | 4 |
| 2020 | Automatic Lyrics Transcription using Dilated Convolutional Neural Networks with Self-AttentionabstractSpeech recognition is a well developed research field so that the current state of the art systems are being used in many applications in the software industry, yet as by today, there still does not exist such robust system for the recognition of words and sentences from singing voice. This paper proposes a complete pipeline for this task which may commonly be referred as automatic lyrics transcription (ALT). We have trained convolutional time-delay neural networks with self-attention on monophonic karaoke recordings using a sequence classification objective for building the acoustic model. The dataset used in this study, DAMP - Sing! 300x30x2 [1] is filtered to have songs with only English lyrics. Different language models are tested including MaxEnt and Recurrent Neural Networks based methods which are trained on the lyrics of pop songs in English. An in-depth analysis of the self-attention mechanism is held while tuning its context width and the number of attention heads. Using the best settings, our system achieves notable improvement to the state-of-the-art in ALT and provides a new baseline for the task. Emir Demirel, Sven Ahlbäck, Simon Dixon |
IJCNN | 3 |
| 2020 | Reliable Local Explanations for Machine ListeningabstractOne way to analyse the behaviour of machine learning models is through local explanations that highlight input features that maximally influence model predictions. Sensitivity analysis, which involves analysing the effect of input perturbations on model predictions, is one of the methods to generate local explanations. Meaningful input perturbations are essential for generating reliable explanations, but there exists limited work on what such perturbations are and how to perform them. This work investigates these questions in the context of machine listening models that analyse audio. Specifically, we use a state-of-the-art deep singing voice detection (SVD) model to analyse whether explanations from SoundLIME (a local explanation method) are sensitive to how the method perturbs model inputs. The results demonstrate that SoundLIME explanations are sensitive to the content in the occluded input regions. We further propose and demonstrate a novel method for quantitatively identifying suitable content type(s) for reliably occluding inputs of machine listening models. The results for the SVD model suggest that the average magnitude of input mel-spectrogram bins is the most suitable content type for temporal explanations. Saumitra Mishra, Emmanouil Benetos, Bob L. T. Sturm, Simon Dixon |
IJCNN | 4 |
| 2019 | Understanding Intonation Trajectories and Patterns of Vocal Notes
Jiajie Dai, Simon Dixon |
MMM (2) | 2 |
| 2018 | Similarity Measures for Vocal-Based Drum Sample Retrieval Using Deep Convolutional Auto-EncodersabstractThe expressive nature of the voice provides a powerful medium for communicating sonic ideas, motivating recent research on methods for query by vocalisation. Meanwhile, deep learning methods have demonstrated state-of-the-art results for matching vocal imitations to imitated sounds, yet little is known about how well learned features represent the perceptual similarity between vocalisations and queried sounds. In this paper, we address this question using similarity ratings between vocal imitations and imitated drum sounds. We use a linear mixed effect regression model to show how features learned by convolutional auto-encoders (CAEs) perform as predictors for perceptual similarity between sounds. Our experiments show that CAEs outperform three baseline feature sets (spectrogram-based representations, MFCCs, and temporal features) at predicting the subjective similarity ratings. We also investigate how the size and shape of the encoded layer effects the predictive power of the learned features. The results show that preservation of temporal information is more important than spectral resolution for this application. Adib Mehrabi, Keunwoo Choi, Simon Dixon, Mark B. Sandler |
ICASSP | 3 |
| 2018 | Towards Complete Polyphonic Music Transcription: Integrating Multi-Pitch Detection and Rhythm QuantizationabstractMost work on automatic transcription produces “piano roll” data with no musical interpretation of the rhythm or pitches. We present a polyphonic transcription method that converts a music audio signal into a human-readable musical score, by integrating multi-pitch detection and rhythm quantization methods. This integration is made difficult by the fact that the multi-pitch detection produces erroneous notes such as extra notes and introduces timing errors that are added to temporal deviations due to musical expression. Thus, we propose a rhythm quantization method that can remove extra notes by extending the metrical hidden Markov model and optimize the model parameters. We also improve the note-tracking process of multi-pitch detection by refining the treatment of repeated notes and adjustment of onset times. Finally, we propose evaluation measures for transcribed scores. Systematic evaluations on commonly used classical piano data show that these treatments improve the performance of transcription, which can be used as benchmarks for further studies. Eita Nakamura, Emmanouil Benetos, Kazuyoshi Yoshii, Simon Dixon |
ICASSP | 4 |
| 2018 | Adversarial Semi-Supervised Audio Source Separation Applied to Singing Voice ExtractionabstractThe state of the art in music source separation employs neural networks trained in a supervised fashion on multi-track databases to estimate the sources from a given mixture. With only few datasets available, often extensive data augmentation is used to combat overfitting. Mixing random tracks, however, can even reduce separation performance as instruments in real music are strongly correlated. The key concept in our approach is that source estimates of an optimal separator should be indistinguishable from real source signals. Based on this idea, we drive the separator towards outputs deemed as realistic by discriminator networks that are trained to tell apart real from separator samples. This way, we can also use unpaired source and mixture recordings without the drawbacks of creating unrealistic music mixtures. Our framework is widely applicable as it does not assume a specific network architecture or number of sources. To our knowledge, this is the first adoption of adversarial training for music source separation. In a prototype experiment for singing voice separation, separation performance increases with our approach compared to purely supervised training. Daniel Stoller, Sebastian Ewert, Simon Dixon |
ICASSP | 3 |
| 2017 | Pickup position and plucking point estimation on an electric guitarabstractThis paper describes a technique to estimate the plucking point and magnetic pickup location along the strings of an electric guitar from a recording of an isolated guitar tone. The estimated values are calculated by minimising the difference between the magnitude spectrum of the recorded tone and that of an electric guitar model based on an ideal string. The recorded tones that are used for the experiment consist of a direct input electric guitar played on all six open strings and played moderately loud (mezzo-forte). The technique is able to estimate pickup locations with 7.75-9.44 mm average absolute error and plucking points with 10.45-10.97 mm average absolute error for single and mixed pickups. Zulfadhli Mohamad, Simon Dixon, Christopher Harte |
ICASSP | 2 |
| 2017 | Towards the characterization of singing styles in world musicabstractIn this paper we focus on the characterization of singing styles in world music.We develop a set of contour features capturing pitch structure and melodic embellishments.Using these features we train a binary classifier to distinguish vocal from non-vocal contours and learn a dictionary of singing style elements.Each contour is mapped to the dictionary elements and each recording is summarized as the histogram of its contour mappings.We use K-means clustering on the recording representations as a proxy for singing style similarity.We observe clusters distinguished by characteristic uses of singing techniques such as vibrato and melisma.Recordings that are clustered together are often from neighbouring countries or exhibit aspects of language and cultural proximity.Studying singing particularities in this comparative manner can contribute to understanding the interaction and exchange between world music styles. Maria Panteli, Rachel M. Bittner, Juan Pablo Bello, Simon Dixon |
ICASSP | 4 |
| 2017 | Tracking metrical structure changes with sparse-NMFabstractThe estimation of rhythmic properties such as tempo, beat positions or metrical structure are central aspects of Music Information Retrieval (MIR) research. Meter inference algorithms are typically designed to track metrical structure in presence of mild deviations of the feature estimates over time in order to account for performance imprecisions, expressive timing or musical effects such as accelerando. Abrupt changes of metrical structure over time are comparatively rarely addressed. In this paper, we present an unsupervised approach to detect metrical structure changes. Formulating the problem as a metrical structure based segmentation retrieval task, we present a variant of sparse NMF and compare it to existing methods. For evaluation, we introduce a new dataset of music recordings containing metric modulations with the corresponding annotations. Elio Quinton, Ken O'Hanlon, Simon Dixon, Mark B. Sandler |
ICASSP | 3 |
| 2017 | A Data-Driven Model of Tonal Chord Sequence ComplexityabstractWe present a compound language model of tonal chord sequences, and evaluate its capability to estimate perceived harmonic complexity. In order to build the compound model, we trained three different models: prediction by partial matching, a hidden Markov model and a deep recurrent neural network on a novel large dataset containing half a million annotated chord sequences. We describe the training process and propose an interpretation of the harmonic patterns that are learned by the hidden states of these models. We use the compound model to generate new chord sequences and estimate their probability, which we then relate to perceived harmonic complexity. In order to collect subjective ratings of complexity, we devised a listening test comprising two different experiments. In the first, subjects choose the more complex chord sequence between two. In the second, subjects rate with a continuous scale the complexity of a single chord sequence. The results of both experiments show a strong relation between negative log probability, given by our language model, and the perceived complexity ratings. The relation is stronger for subjects with high musical sophistication index, acquired through the GoldMSI standard questionnaire. The analysis of the results also includes the preference ratings that have been collected along with the complexity ratings; a weak negative correlation emerged between preference and log probability. Bruno Di Giorgi, Simon Dixon, Massimiliano Zanoni, Augusto Sarti |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Note Value Recognition for Piano Transcription Using Markov Random FieldsabstractThis paper presents a statistical method for use in music transcription that can estimate score times of note onsets and offsets from polyphonic MIDI performance signals. Because performed note durations can deviate largely from score-indicated values, previous methods had the problem of not being able to accurately estimate offset score times (or note values) and, thus, could only output incomplete musical scores. Based on observations that the pitch context and onset score times are influential on the configuration of note values, we construct a context-tree model that provides prior distributions of note values using these features and combine it with a performance model in the framework of Markov random fields. Evaluation results show that our method reduces the average error rate by around 40 percent compared to existing/simple methods. We also confirmed that, in our model, the score model plays a more important role than the performance model, and it automatically captures the voice structure by unsupervised learning. Eita Nakamura, Kazuyoshi Yoshii, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Identifying Missing and Extra Notes in Piano Recordings Using Score-Informed Dictionary LearningabstractThe goal of automatic music transcription (AMT) is to obtain a high-level symbolic representation of the notes played in a given audio recording. Despite being researched for several decades, current methods are still inadequate for many applications. To boost the accuracy in a music tutoring scenario, we exploit that the score to be played is specified and we only need to detect the differences to the actual performance. In contrast to previous work that uses score information for postprocessing, we employ the score to construct a transcription method that is tailored to the given audio recording. By adapting a score-informed dictionary learning technique as used for source separation, we learn for each score pitch a spectral pattern describing the energy distribution of associated notes in the recording. In this paper, we identify several systematic weaknesses in our previous approach and introduce three extensions to improve its performance. First, we extend our dictionary of spectral templates to a dictionary of variable-length spectrotemporal patterns. Second, we integrate the score information using soft rather than hard constraints, to better take into account that differences from the score indeed occur. Third, we introduce new regularizers to guide the learning process. Our experiments show that these extensions particularly improve the accuracy for identifying extra notes, while the accuracy for correct and missing notes remains at a similar level. The influence of each extension is demonstrated with further experiments. Siying Wang 0001, Sebastian Ewert, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Estimation of the reliability of multiple rhythm features extraction from a single descriptorabstractThe design of systems for automatic audio feature extraction is a central aspect of the field of Music Information Retrieval. However, feature extraction systems often do not provide an indication of the reliability of the corresponding feature. Nevertheless, the provision of a reliability or confidence measure can be critical for the usage of a given feature in complex systems and real-world applications. In the present study we investigate the relationship between the entropy of a rhythmogram, which has been proposed as a descriptor of tempo salience in previous work, and the reliability of the extraction of multiple high level rhythm related features. The results show that this single descriptor is viable for simultaneously estimating the reliability of multiple rhythm features extraction. The results also provide quantitative insight that is consistent with qualitative observations extensively reported in the literature on a qualitative basis. Elio Quinton, Mark B. Sandler, Simon Dixon |
ICASSP | 3 |
| 2016 | An End-to-End Neural Network for Polyphonic Piano Music TranscriptionabstractWe present a supervised neural network model for polyphonic piano music transcription. The architecture of the proposed model is analogous to speech recognition systems and comprises an acoustic model and a music language model. The acoustic model is a neural network used for estimating the probabilities of pitches in a frame of audio. The language model is a recurrent neural network that models the correlations between pitch combinations over time. The proposed model is general and can be used to transcribe polyphonic music without imposing any constraints on the polyphony. The acoustic and language model predictions are combined using a probabilistic graphical model. Inference over the output variables is performed using the beam search algorithm. We perform two sets of experiments. We investigate various neural network architectures for the acoustic models and also investigate the effect of combining acoustic and music language model predictions using the proposed architecture. We compare performance of the neural network-based acoustic models with two popular unsupervised acoustic models. Results show that convolutional neural network acoustic models yield the best performance across all evaluation metrics. We also observe improved performance with the application of the music language models. Finally, we present an efficient variant of beam search that improves performance and reduces run-times by an order of magnitude, making the model suitable for real-time applications. Siddharth Sigtia, Emmanouil Benetos, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | Robust and Efficient Joint Alignment of Multiple Musical PerformancesabstractThe goal of music alignment is to map each temporal position in one version of a piece of music to the corresponding positions in other versions of the same piece. Despite considerable improvements in recent years, state-of-the-art methods still often fail to identify a correct alignment if versions differ substantially with respect to acoustic conditions or musical interpretation. To increase the robustness for these cases, we exploit in this work the availability of multiple versions of the piece to be aligned. By processing these jointly, we can supply the alignment process with additional examples of how a section might be interpreted or which acoustic conditions may arise. This way, we can use alignment information between two versions transitively to stabilize the alignment with a third version. Extending our previous work [1], we present two such joint alignment methods, progressive alignment and probabilistic profile, and discuss their fundamental differences and similarities on an algorithmic level. Our systematic experiments using 376 recordings of 9 pieces demonstrate that both methods can indeed improve the alignment accuracy and robustness over comparable pairwise methods. Further, we provide an in-depth analysis of the behavior of both joint alignment methods, studying the influence of parameters such as the number of performances available, comparing their computational costs, and investigating further strategies to increase both. Siying Wang 0001, Sebastian Ewert, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2015 | An Evaluation of Multidimensional Controllers for Sound Design TasksabstractThis paper presents an investigation into musicians' ability to control sound synthesiser parameters using various inter- faces. The principal aim was to compare separate, 1D parameter controls (touchscreen sliders) to multidimensional con- trollers (an XY touchpad for 2D, the Leap Motion for 3D). Subjects had to match a target sound as quickly and accurately as possible. Results show that after about two hours of practice, the XY pad is 9% faster than two sliders for no accuracy loss, and the Leap is 17% faster than 3 sliders with 9% accuracy loss. The multidimensional controllers improved most with practice. A new perspective on Fitts' index of difficulty is presented: "Index of Search Space Reduction" (ISSR). ISSR and retrospective accuracy thresholds on the search trajectory are used to obtain straight line plots and throughput values. These plots reveal that the Leap's speed improvement was mainly due to reaction time, but the XY pad traversed the space faster. Robert Tubb, Simon Dixon |
CHI | 2 |
| 2015 | Modelling the decay of piano soundsabstractWe investigate piano acoustics and compare the theoretical temporal decay of individual partials to recordings of real-world piano notes from the RWC Music Database. We first describe the theory behind double decay and beats, known phenomena caused by the interaction between strings and soundboard. Then we fit the decay of the first 30 partials to a standard linear model and two physically-motivated non-linear models that take into account the coupling of strings and soundboard. We show that the use of non-linear models provides a better fit to the data. We use these estimated decay rates to parameterise the characteristic decay response (decay rates along frequencies) of the piano under investigation. The results also show that dynamics have no significant effect on the decay rate. Tian Cheng 0001, Simon Dixon, Matthias Mauch |
ICASSP | 2 |
| 2015 | A hybrid recurrent neural network for music transcriptionabstractWe investigate the problem of incorporating higher-level symbolic score-like information into Automatic Music Transcription (AMT) systems to improve their performance. We use recurrent neural networks (RNNs) and their variants as music language models (MLMs) and present a generative architecture for combining these models with predictions from a frame level acoustic classifier. We also compare different neural network architectures for acoustic modeling. The proposed model computes a distribution over possible output sequences given the acoustic input signal and we present an algorithm for performing a global search for good candidate transcriptions. The performance of the proposed model is evaluated on piano music from the MAPS dataset and we observe that the proposed model consistently outperforms existing transcription methods. Siddharth Sigtia, Emmanouil Benetos, Nicolas Boulanger-Lewandowski, Tillman Weyde, Artur S. d'Avila Garcez, Simon Dixon |
ICASSP | 6 |
| 2015 | Compensating for asynchronies between musical voices in score-performance alignmentabstractThe goal of score-performance synchronisation is to align a given musical score to an audio recording of a performance of the same piece. A major challenge in computing such alignments is to account for musical parameters including the local tempo or playing style. To increase the overall robustness, current methods assume that notes occurring simultaneously in the score are played concurrently in a performance. Musical voices such as the melody, however, are often played asynchronously to other voices, which can lead to significant local alignment errors. In this paper, we present a novel method that handles asynchronies between the melody and the accompaniment by treating the voices as separate time lines in a multi-dimensional variant of dynamic time warping (DTW). Constraining the alignment with information obtained via classical DTW, our method measurably improves the alignment accuracy for pieces with asynchronous voices and preserves the accuracy otherwise. Siying Wang 0001, Sebastian Ewert, Simon Dixon |
ICASSP | 3 |
| 2015 | Identifying Cover Songs Using Information-Theoretic Measures of SimilarityabstractThis paper investigates methods for quantifying similarity between audio signals, specifically for the task of cover song detection. We consider an information-theoretic approach, where we compute pairwise measures of predictability between time series. We compare discrete-valued approaches operating on quantized audio features, to continuous-valued approaches. In the discrete case, we propose a method for computing the normalized compression distance, where we account for correlation between time series. In the continuous case, we propose to compute information-based measures of similarity as statistics of the prediction error between time series. We evaluate our methods on two cover song identification tasks using a data set comprised of 300 Jazz standards and using the Million Song Dataset. For both datasets, we observe that continuous-valued approaches outperform discrete-valued approaches. We consider approaches to estimating the normalized compression distance (NCD) based on string compression and prediction, where we observe that our proposed normalized compression distance with alignment (NCDA) improves average performance over NCD, for sequential compression algorithms. Finally, we demonstrate that continuous-valued distances may be combined to improve performance with respect to baseline approaches. Using a large-scale filter-and-refine approach, we demonstrate state-of-the-art performance for cover song identification using the Million Song Dataset. Peter Foster, Simon Dixon, Anssi Klapuri |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | A Comparison of Extended Source-Filter Models for Musical Signal Reconstruction
Tian Cheng 0001, Simon Dixon, Matthias Mauch |
DAFx | 2 |
| 2014 | Towards complex matrix decomposition of spectrograms based on the relative phase offsets of harmonic soundsabstractIn this paper we study the relative phase offsets between partials in the sustained part of harmonic sounds and investigate their suitability for complex matrix decomposition of spectrograms. We formally introduce this property in a sinusoidal model and visualise the phase relations of a musical instrument. A model of complex matrix decomposition in the time-frequency domain is derived and equations for the estimation of the model parameters are provided in the mono-phonic case. We illustrate the model with the analysis of a mono-phonic saxophone signal. The results suggest that the phase offset is able to capture inherent time-invariant phase properties of harmonic sounds and outline its potential use for complex matrix decomposition. Holger Kirchhoff, Roland Badeau, Simon Dixon |
ICASSP | 3 |
| 2014 | PYIN: A fundamental frequency estimator using probabilistic threshold distributionsabstractWe propose the Probabilistic YIN (PYIN) algorithm, a modification of the well-known YIN algorithm for fundamental frequency (F0) estimation. Conventional YIN is a simple yet effective method for frame-wise monophonic F0 estimation and remains one of the most popular methods in this domain. In order to eliminate short-term errors, outputs of frequency estimators are usually post-processed resulting in a smoother pitch track. One shortcoming of YIN is that such post-processing cannot fall back on alternative interpretations of the signal because the method outputs precisely one estimate per frame. To address this problem we modify YIN to output multiple pitch candidates with associated probabilities (PYIN Stage 1). These probabilities arise naturally from a prior distribution on the YIN threshold parameter. We use these probabilities as observations in a hidden Markov model, which is Viterbi-decoded to produce an improved pitch track (PYIN Stage 2). We demonstrate that the combination of Stages 1 and 2 raises recall and precision substantially. The additional computational complexity of PYIN over YIN is low. We make the method freely available online1as an open source C++ library for Vamp hosts. Matthias Mauch, Simon Dixon |
ICASSP | 2 |
| 2014 | Improved music feature learning with deep neural networksabstractRecent advances in neural network training provide a way to efficiently learn representations from raw data. Good representations are an important requirement for Music Information Retrieval (MIR) tasks to be performed successfully. However, a major problem with neural networks is that training time becomes prohibitive for very large datasets and the learning algorithm can get stuck in local minima for very deep and wide network architectures. In this paper we examine 3 ways to improve feature learning for audio data using neural networks: 1.using Rectified Linear Units (ReLUs) instead of standard sigmoid units; 2.using a powerful regularisation technique called Dropout; 3.using Hessian-Free (HF) optimisation to improve training of sigmoid nets. We show that these methods provide significant improvements in training time and the features learnt are better than state of the art handcrafted features, with a genre classification accuracy of 83 ± 1.1% on the Tzanetakis (GTZAN) dataset. We found that the rectifier networks learnt better features than the sigmoid networks. We also demonstrate the capacity of the features to capture relevant information from audio data by applying them to genre classification on the ISMIR 2004 dataset. Siddharth Sigtia, Simon Dixon |
ICASSP | 2 |
| 2014 | A Four Strategy Model of Creative Parameter Space Interaction
Robert Tubb, Simon Dixon |
ICCC | 2 |
| 2014 | Sequential complexity as a descriptor for musical similarityabstractWe propose string compressibility as a descriptor of temporal structure in audio, for the purpose of determining musical similarity. Our descriptors are based on computing track-wise compression rates of quantized audio features, using multiple temporal resolutions and quantization granularities. To verify that our descriptors capture musically relevant information, we incorporate our descriptors into similarity rating prediction and song year prediction tasks. We base our evaluation on a dataset of 15 500 track excerpts of Western popular music, for which we obtain 7 800 web-sourced pairwise similarity ratings. To assess the agreement among similarity ratings, we perform an evaluation under controlled conditions, obtaining a rank correlation of 0.33 between intersected sets of ratings. Combined with bag-of-features descriptors, we obtain performance gains of 31.1% and 10.9% for similarity rating prediction and song year prediction. For both tasks, analysis of selected descriptors reveals that representing features at multiple time scales benefits prediction accuracy. Peter Foster, Matthias Mauch, Simon Dixon |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Identification of cover songs using information theoretic measures of similarityabstractWe consider techniques for cover song detection, based on information theoretic notions of compressibility. We propose methods for computing the normalised compression distance (NCD), while accounting for correlation between time series. Secondly, we describe methods based on cross-prediction for estimating compressibility between sequences of continuous-valued features. Using the latter approach, we view the NCD as a statistic of the prediction error. We evaluate the proposed approaches using a data set consisting of 300 Jazz songs. Quantified in terms of mean average precision, the proposed continuous-valued approach outperforms considered quantisation-based approaches. Peter Foster, Simon Dixon, Anssi Klapuri |
ICASSP | 2 |
| 2013 | Missing template estimation for user-assisted music transcriptionabstractFor a user-assisted music transcription system in which the user is asked to label some notes for each instrument in the recording, we investigate ways to limit the amount of information the user has to provide. Different methods are proposed and experimentally compared that enable the estimation of template spectra at pitch positions that have not been annotated by the user, in order to derive a full set of instrument templates that can be used within a non-negative matrix factorisation framework. A set of error metrics is presented that enables the evaluation of the NMF gain matrix. The results show that purely data-driven methods outperform more refined instrument models when the user annotates notes at many different pitches for each instrument. When notes are labelled at a smaller number of different pitches, the highest accuracies are obtained using pre-stored instrument templates that are adapted to the instruments in the mixture. Holger Kirchhoff, Simon Dixon, Anssi Klapuri |
ICASSP | 2 |
| 2013 | Automatic music transcription: challenges and future directions
Emmanouil Benetos, Simon Dixon, Dimitrios Giannoulis, Holger Kirchhoff, Anssi Klapuri |
J. Intell. Inf. Syst. | 2 |
| 2012 | Shift-variant non-negative matrix deconvolution for music transcriptionabstractIn this paper, we address the task of semi-automatic music transcription in which the user provides prior information about the polyphonic mixture under analysis. We propose a non-negative matrix deconvolution framework for this task that allows instruments to be represented by a different basis function for each fundamental frequency (“shift variance”). Two different types of user input are studied: information about the types of instruments, which enables the use of basis functions from an instrument database, and a manual transcription of a number of notes which enables the template estimation from the data under analysis itself. Experiments are performed on a data set of mixtures of acoustical instruments up to a polyphony of five. The results confirm a significant loss in accuracy when database templates are used and show the superiority of the Kullback-Leibler divergence over the least squares error cost function. Holger Kirchhoff, Simon Dixon, Anssi Klapuri |
ICASSP | 2 |
| 2011 | Polyphonic music transcription using note onset and offset detectionabstractIn this paper, an approach for polyphonic music transcription based on joint multiple-F0 estimation and note onset/offset detection is proposed. For preprocessing, the resonator time-frequency image of the input music signal is extracted and noise suppression is performed. A pitch salience function is extracted for each frame along with tuning and inharmonicity parameters. For onset detection, late fusion is employed by combining a novel spectral flux-based feature which incorporates pitch tuning information and a novel salience function-based descriptor. For each segment defined by two onsets, an overlapping partial treatment procedure is used and a pitch set score function is proposed. A note offset detection procedure is also proposed using HMMs trained on MIDI data. The system was trained on piano chords and tested on classic and jazz recordings from the RWC database. Improved transcription results are reported compared to state-of-the-art approaches. Emmanouil Benetos, Simon Dixon |
ICASSP | 2 |
| 2011 | Real-time synchronisation of multimedia streams in a mobile deviceabstractWith the constant improvements in the technical capabilities and bandwidth available to mobile phones, mobile audio and video streaming services are booming. This allows music enthusiasts to watch their favourite music videos instead of only listening to the audio tracks, thus augmenting their listening experience to be multimodal. Usually the highly compressed audio tracks in these videos result in poorer quality music, compared to music the user might already have locally on their phone or playing through another sound source. In this paper we present MuViSync Mobile, a mobile phone application that synchronises real-time high quality music, either stored locally or input through the microphone, with the corresponding streaming music video. We extend previous work on music to music video synchronisation by proposing an alternative algorithm for higher efficiency with similar alignment accuracy, which we tested with music video examples including simulated noise. This algorithm correctly aligns 90% of the audio frames to within 100 ms of the known alignment. We also describe its implementation on an iPhone. Robert Macrae, Joachim Neumann, Xavier Anguera Miró, Nuria Oliver, Simon Dixon |
ICME | 5 |
| 2010 | High precision frequency estimation for harpsichord tuning classificationabstractWe present a novel music signal processing task of classifying the tuning of a harpsichord from audio recordings of standard musical works. We report the results of a classification experiment involving six different temperaments, using real harpsichord recordings as well as synthesised audio data. We introduce the concept of conservative transcription, and show that existing high-precision pitch estimation techniques are sufficient for our task if combined with conservative transcription. In particular, using the CQIFFT algorithm with conservative transcription and removal of short duration notes, we are able to distinguish between 6 different temperaments of harpsichord recordings with 96% accuracy (100% for synthetic data). Dan Tidhar, Matthias Mauch, Simon Dixon |
ICASSP | 3 |
| 2010 | A guitar tablature score followerabstractAlthough guitar tablature is the most prevalent musical score format on the internet, score following programs are only made for traditional musical scores or MIDI. The following work demonstrates a score following method for guitar tablature (tabs). In this method, a guitar tab parser interprets the score from HTML or ASCII text tabs which is then synthesised with guitar-like parameters. Finally, real-time audio to audio synchronisation methods are then used to accurately track the musicians position within the tab. A prototype application has been made to demonstrate this technique. Robert Macrae, Simon Dixon |
ICME | 2 |
| 2010 | Simultaneous Estimation of Chords and Musical Context From AudioabstractChord labels provide a concise description of musical harmony. In pop and jazz music, a sequence of chord labels is often the only written record of a song, and forms the basis of so-called lead sheets. We devise a fully automatic method to simultaneously estimate from an audio waveform the chord sequence including bass notes, the metric positions of chords, and the key. The core of the method is a six-layered dynamic Bayesian network, in which the four hidden source layers jointly model metric position, key, chord, and bass pitch class, while the two observed layers model low-level audio features corresponding to bass and treble tonal content. Using 109 different chords our method provides substantially more harmonic detail than previous approaches while maintaining a high level of accuracy. We show that with 71% correctly classified chords our method significantly exceeds the state of the art when tested against manually annotated ground truth transcriptions on the 176 audio tracks from the MIREX 2008 Chord Detection Task. We introduce a measure of segmentation quality and show that bass and meter modeling are especially beneficial for obtaining the correct level of granularity. Matthias Mauch, Simon Dixon |
IEEE Trans. Speech Audio Process. | 2 |
| 2008 | Automatic Page Turning for Musicians via Real-Time Machine ListeningabstractWe present a system that automatically turns the pages of the music score for musicians during a performance. It is based on a new algorithm for following an incoming audio stream in real time and aligning it to a music score (in the form of a synthesised audio file). Precision and robustness of the algorithm are quantified in systematic experiments, and a demonstration using an actual page turning machine built by an Austrian company is described. Andreas Arzt, Gerhard Widmer, Simon Dixon |
ECAI | 3 |
| 2007 | Evaluating Low-Level Features for Beat Classification and TrackingabstractIn this paper, we address the question of which low-level acoustical features are the most suitable for identifying music beats computationally. We consider 172 features computed on consecutive signal frames and systematically evaluate their individual value in the task of providing reliable cues for the presence and localisation of beats in music signals. We compare two ways of evaluating features: their accuracy in a song-specific classification task (classifying beats vs nonbeats) and their performance as a front-end to a beat tracking system. Fabien Gouyon, Simon Dixon, Gerhard Widmer |
ICASSP (4) | 2 |
| 2006 | An experimental comparison of audio tempo induction algorithmsabstractWe report on the tempo induction contest organized during the International Conference on Music Information Retrieval (ISMIR 2004) held at the University Pompeu Fabra in Barcelona, Spain, in October 2004. The goal of this contest was to evaluate some state-of-the-art algorithms in the task of inducing the basic tempo (as a scalar, in beats per minute) from musical audio signals. To our knowledge, this is the first published large scale cross-validation of audio tempo induction algorithms. Participants were invited to submit algorithms to the contest organizer, in one of several allowed formats. No training data was provided. A total of 12 entries (representing the work of seven research teams) were evaluated, 11 of which are reported in this document. Results on the test set of 3199 instances were returned to the participants before they were made public. Anssi Klapuri's algorithm won the contest. This evaluation shows that tempo induction algorithms can reach over 80% accuracy for music with a constant tempo, if we do not insist on finding a specific metrical level. After the competition, the algorithms and results were analyzed in order to discover general lessons for the future development of tempo induction systems. One conclusion is that robust tempo induction entails the processing of frame features rather than that of onset lists. Further, we propose a new "redundant" approach to tempo induction, inspired by knowledge of human perceptual mechanisms, which combines multiple simpler methods using a voting mechanism. Machine emulation of human tempo induction is still an open issue. Many avenues for future work in audio tempo tracking are highlighted, as for instance the definition of the best rhythmic features and the most appropriate periodicity detection method. In order to stimulate further research, the contest results, annotations, evaluation software and part of the data are available at http://ismir2004.ismir.net/ISMIR_Contest.html Fabien Gouyon, Anssi Klapuri, Simon Dixon, M. Alonso, George Tzanetakis, C. Uhle, Pedro Cano |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | An On-Line Time Warping Algorithm for Tracking Musical Performances
Simon Dixon |
IJCAI | 1 |
| 2000 | Beat Tracking with Musical Knowledge
Simon Dixon, Emilios Cambouropoulos |
ECAI | 1 |
| 2000 | A Lightweight Multi-agent Musical Beat Tracking System
Simon Dixon |
PRICAI | 1 |
| 1993 | The Implementation of a First-Order Logic AGM Belief Revision SystemabstractBelief revision is increasingly being seen as central to a number of fundamental problems in artificial intelligence such as nonmonotonic reasoning, reasoning about action, truth maintenance and database update. The authors describe the first implementation of an AGM belief revision system. The system is based on classical first-order logic, and for any finitely representable belief state, it efficiently computes expansions, contractions and revision satisfying the AGM postulates for rational belief change. The system uses a finite base to represent a belief set, and interprets a partially specified entrenchment as representing a unique most conservative entrenchment-this is motivated by considerations of evidence and by the close connections between belief revision and nonmonotonic reasoning. The authors describe in detail the algorithms for belief change, and give some examples of the system's operation. Simon Dixon, Wayne Wobcke |
ICTAI | 1 |
| 1993 | Connections Between the ATMS and AGM Belief Revision
Simon Dixon, Norman Y. Foo |
IJCAI | 1 |