György Fazekas

dblp:13/9955 · also George Fazekas · DBLP profile ↗
← Back
41ranked-venue papers
1as first author
21since 2021 · last 2026
0000-0003-2580-0007ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 11 since 2021Artificial intelligence and machine learning · 11 · 7 since 2021Databases, data management, data science and information retrieval · 10 · 4 since 2021Computer networks · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Visualising Pianists' Touch: Transcribing Expressive Piano Performance from Audio to Piano Key Motion
abstract
Detailed measurements of piano key motion capture touch, timing, and dynamic control, providing crucial performance insights. Such expressive gestures are overlooked in MIDI, which only records pitch onset, duration, and velocity. Here, we introduce a novel transcription technique that directly maps audio from expressive piano performance to continuous piano key motion. User studies reveal a preference to the transcribed key motion trajectories over MIDI in representing sound, and over 80% accuracy in matching transcribed trajectories to audio from contrasting piano expressions. Follow-up interviews further indicate that the visualised trajectories can reveal subtle performance nuances and provide actionable guidance for both teaching and practice. An interface example for pedagogy and performance analysis utilising our technique is also illustrated. By providing a physically grounded performance representation that musicians can interpret and act upon, this work establishes a foundation for future interactive tools in music pedagogy, performance feedback, and embodied musical learning.
Jingjing Tang 0002, Shinichi Furuya, Hayato Nishioka, Momoko Shioki, Geraint A. Wiggins, György Fazekas, Vincent K. M. Cheung
CHI6
2025 Leave-One-EquiVariant: Alleviating Invariance-Related Information Loss in Contrastive Music Representations
abstract
Contrastive learning has proven effective in self-supervised musical representation learning, particularly for Music Information Retrieval (MIR) tasks. However, reliance on augmentation chains for contrastive view generation and the resulting learnt invariances pose challenges when different downstream tasks require sensitivity to certain musical attributes. To address this, we propose the Leave One EquiVariant (LOEV) framework, which introduces a flexible, task-adaptive approach compared to previous work by selectively preserving information about specific augmentations, allowing the model to maintain task-relevant equivariances. We demonstrate that LOEV alleviates information loss related to learned invariances, improving performance on augmentation related tasks and retrieval without sacrificing general representation quality. Furthermore, we introduce a variant of LOEV, LOEV++, which builds a disentangled latent space by design in a self-supervised manner, and enables targeted retrieval based on augmentation related attributes.
Julien Guinot, Elio Quinton, György Fazekas
ICASSP3
2025 Music2Latent2: Audio Compression with Summary Embeddings and Autoregressive Decoding
abstract
Efficiently compressing high-dimensional audio signals into a compact and informative latent space is crucial for various tasks, including generative modeling and music information retrieval (MIR). Existing audio autoencoders, however, often struggle to achieve high compression ratios while preserving audio fidelity and facilitating efficient downstream applications. We introduce Music2Latent2, a novel audio autoencoder that addresses these limitations by leveraging consistency models and a novel approach to representation learning based on unordered latent embeddings, which we call summary embeddings. Unlike conventional methods that encode local audio features into ordered sequences, Music2Latent2 compresses audio signals into sets of summary embeddings, where each embedding can capture distinct global features of the input sample. This enables to achieve higher reconstruction quality at the same compression ratio. To handle arbitrary audio lengths, Music2Latent2 employs an autoregressive consistency model trained on two consecutive audio chunks with causal masking, ensuring coherent reconstruction across segment boundaries. Additionally, we propose a novel two-step decoding procedure that leverages the denoising capabilities of consistency models to further refine the generated audio at no additional cost. Our experiments demonstrate that Music2Latent2 outperforms existing continuous audio autoencoders regarding audio quality and performance on downstream tasks. Music2Latent2 paves the way for new possibilities in audio compression.
Marco Pasini, Stefan Lattner, György Fazekas
ICASSP3
2025 Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
abstract
This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performances from score inputs. To the best of our knowledge, this is the first system to offer a streamlined method for converting score MIDI files lacking expression control into rich, expressive piano performances. We conducted experiments using subsets of the ATEPP dataset, evaluating the system with both objective metrics and subjective listening tests. Our system not only accurately reconstructs human-like expressiveness, but also captures the acoustic ambience of environments such as concert halls and recording studios. Additionally, the proposed system demonstrates its ability to achieve musical expressiveness while ensuring good audio quality in its outputs.
Jingjing Tang 0002, Erica Cooper, Xin Wang 0037, Junichi Yamagishi, György Fazekas
ICASSP5
2025 Position Paper: Towards a Unified Representation Evaluation Framework Beyond Downstream Tasks
abstract
Downstream probing has been the dominant method for evaluating model representations, an important process given the increasing prominence of self-supervised learning and foundation models. However, downstream probing primarily assesses the availability of task-relevant information in the model’s latent space, overlooking attributes such as equivariance, invariance, and disentanglement, which contribute to the interpretability, adaptability, and utility of representations in real-world applications. While some attempts have been made to measure these qualities in representations, no unified evaluation framework with modular, generalizable, and interpretable metrics exists.In this paper, we argue for the importance of representation evaluation beyond downstream probing. We introduce a standardized protocol to quantify informativeness, equivariance, invariance, and disentanglement of factors of variation in model representations. We use it to evaluate representations from a variety of models in the image and speech domains using different architectures and pretraining approaches on identified controllable factors of variation. We find that representations from models with similar downstream performance can behave substantially differently with regard to these attributes. This hints that the respective mechanisms underlying their downstream performance are functionally different, prompting new research directions to understand and improve representations.
Christos Plachouras, Julien Guinot, György Fazekas, Elio Quinton, Emmanouil Benetos, Johan Pauwels
IJCNN3
2025 Mamba-Diffusion Model with Learnable Wavelet for Controllable Symbolic Music Generation
abstract
The recent surge in the popularity of diffusion models for image synthesis has attracted new attention to their potential for generation tasks in other domains. However, their applications to symbolic music generation remain largely under-explored because symbolic music is typically represented as sequences of discrete events and standard diffusion models are not well-suited for discrete data. We represent symbolic music as image-like pianorolls, facilitating the use of diffusion models for the generation of symbolic music. Moreover, this study introduces a novel diffusion model that incorporates our proposed Transformer-Mamba block and learnable wavelet transform. Classifier-free guidance is utilised to generate symbolic music with target chords. Our evaluation shows that our method achieves compelling results in terms of music quality and controllability, outperforming the strong baseline in pianoroll generation. Our code is available at https://github.com/jinchengzhanggg/proffusion.
György Fazekas, Charalampos Saitis
IJCNN2
2025 Low-Data Classification of Historical Music Manuscripts: A Few-Shot Learning Approach
abstract
In this paper, we explore the intersection of technology and cultural preservation by developing a self-supervised learning framework for the classification of musical symbols in historical manuscripts. Optical Music Recognition (OMR) plays a vital role in digitising and preserving musical heritage, but historical documents often lack the labelled data required by traditional methods. We overcome this challenge by training a neural-based feature extractor on unlabelled data, enabling effective classification with minimal samples. Key contributions include optimising crop preprocessing for a self-supervised Convolutional Neural Network and evaluating classification methods, including SVM, multilayer perceptrons, and prototypical networks. Our experiments yield an accuracy of 87.66%, showcasing the potential of AI-driven methods to ensure the survival of historical music for future generations through advanced digital archiving techniques.
Elona Shatri, Daniel Raymond, György Fazekas
IPAS3
2024 Synthesising Handwritten Music with GANs: A Comprehensive Evaluation of CycleWGAN, ProGAN, and DCGAN
abstract
The generation of handwritten music sheets is a crucial step toward enhancing Optical Music Recognition (OMR) systems, which rely on large and diverse datasets for optimal performance. However, handwritten music sheets, often found in archives, present challenges for digitisation due to their fragility, varied handwriting styles, and image quality. This paper addresses the data scarcity problem by applying Generative Adversarial Networks (GANs) to synthesise realistic handwritten music sheets. We provide a comprehensive evaluation of three GAN models—DCGAN, ProGAN, and CycleWGAN—comparing their ability to generate diverse and high-quality handwritten music images. The proposed CycleWGAN model, which enhances style transfer and training stability, significantly outperforms DCGAN and ProGAN in both qualitative and quantitative evaluations. CycleWGAN achieves superior performance, with an FID score of 41.87, an IS of 2.29, and a KID of 0.05, making it a promising solution for improving OMR systems.
Elona Shatri, Kalikidhar Palavala, György Fazekas
IEEE Big Data3
2024 Differentiable Time-Varying Linear Prediction in the Context of End-to-End Analysis-by-Synthesis
Chin-Yun Yu, György Fazekas
INTERSPEECH2
2023 Rigid-Body Sound Synthesis with Differentiable Modal Resonators
abstract
Physical models of rigid bodies are used for sound synthesis in applications from virtual environments to music production. Traditional methods, such as modal synthesis, often rely on computationally expensive numerical solvers, while recent deep learning approaches are limited by post-processing of their results. In this work, we present a novel end-to-end framework for training a deep neural network to generate modal resonators for a given 2D shape and material using a bank of differentiable IIR filters. We demonstrate our method on a dataset of synthetic objects but train our model using an audio-domain objective, paving the way for physically-informed synthesisers to be learned directly from recordings of real-world objects.
Ben Hayes, Charalampos Saitis, György Fazekas, Mark B. Sandler
ICASSP4
2023 Sinusoidal Frequency Estimation by Gradient Descent
abstract
Sinusoidal parameter estimation is a fundamental task in applications from spectral analysis to time-series forecasting. Estimating the sinusoidal frequency parameter by gradient descent is, however, often impossible as the error function is non-convex and densely populated with local minima. The growing family of differentiable signal processing methods has therefore been unable to tune the frequency of oscillatory components, preventing their use in a broad range of applications. This work presents a technique for joint sinusoidal frequency and amplitude estimation using the Wirtinger derivatives of a complex exponential surrogate and any first order gradient-based optimizer, enabling end-to-end training of neural network controllers for unconstrained sinusoidal models.
Ben Hayes, Charalampos Saitis, György Fazekas
ICASSP3
2023 HIPI: A Hierarchical Performer Identification Model Based on Symbolic Representation of Music
abstract
Automatic Performer Identification from the symbolic representation of music has been a challenging topic in Music Information Retrieval (MIR).In this study, we apply a Recurrent Neural Network (RNN) model to classify the most likely music performers from their interpretative styles. We study different expressive parameters and investigate how to quantify these parameters for the exceptionally challenging task of performer identification. We encode performerstyle information using a Hierarchical Attention Network (HAN) architecture, based on the notion that traditional western music has a hierarchical structure (note, beat, measure, phrase level etc.). In addition, we present a large-scale dataset consisting of six virtuoso pianists performing the same set of compositions. The experimental results show that our model outperforms the baseline models with an F1-score of 0.845 and demonstrates the significance of the attention mechanism for understanding different performance styles.
Syed Rifat Mahmud Rafee, György Fazekas, Geraint A. Wiggins
ICASSP2
2023 Conditioning and Sampling in Variational Diffusion Models for Speech Super-Resolution
abstract
Recently, diffusion models (DMs) have been increasingly used in audio processing tasks, including speech super-resolution (SR), which aims to restore high-frequency content given low-resolution speech utterances. This is commonly achieved by conditioning the network of noise predictor with low-resolution audio. In this paper, we propose a novel sampling algorithm that communicates the information of the low-resolution audio via the reverse sampling process of DMs. The proposed method can be a drop-in replacement for the vanilla sampling process and can significantly improve the performance of the existing works. Moreover, by coupling the proposed sampling method with an unconditional DM, i.e., a DM with no auxiliary inputs to its noise predictor, we can generalize it to a wide range of SR setups. We also attain state-of-the-art results on the VCTK Multi-Speaker benchmark with this novel formulation.
Chin-Yun Yu, Sung-Lin Yeh, György Fazekas, Hao Tang 0002
ICASSP3
2023 The Internet of Sounds: Convergent Trends, Insights, and Future Directions
abstract
Current sound-based practices and systems developed in both academia and industry point to convergent research trends that bring together the field of Sound and Music Computing with that of the Internet of Things. This paper proposes a vision for the emerging field of the Internet of Sounds (IoS), which stems from such disciplines. The IoS relates to the network of Sound Things, i.e., devices capable of sensing, acquiring, processing, actuating, and exchanging data serving the purpose of communicating sound-related information. In the IoS paradigm, which merges under a unique umbrella the emerging fields of the Internet of Musical Things and the Internet of Audio Things, heterogeneous devices dedicated to musical and non-musical tasks can interact and cooperate with one another and with other things connected to the Internet to facilitate sound-based services and applications that are globally available to the users. We survey the state of the art in this space, discuss the technological and non-technological challenges ahead of us and propose a comprehensive research agenda for the field.
Luca Turchet, Mathieu Lagrange, Cristina Rottondi, György Fazekas, Nils Peters, Jan Østergaard, Frederic Font, Tom Bäckström, Carlo Fischione
IEEE Internet Things J.4
2023 Semantic integration of audio content providers through the Audio Commons Ontology
Miguel Ceriani, Fabio Viola, Sasa Rudan, Francesco Antoniazzi, Mathieu Barthet, György Fazekas
J. Web Semant.6
2022 Learning Music Audio Representations Via Weak Language Supervision
abstract
Audio representations for music information retrieval are typically learned via supervised learning in a task-specific fashion. Although effective at producing state-of-the-art results, this scheme lacks flexibility with respect to the range of applications a model can have and requires extensively annotated datasets. In this work, we pose the question of whether it may be possible to exploit weakly aligned text as the only supervisory signal to learn general-purpose music audio representations. To address this question, we design a multimodal architecture for music and language pre-training (MuLaP) optimised via a set of proxy tasks. Weak supervision is provided in the form of noisy natural language descriptions conveying the overall musical content of the track. After pre-training, we transfer the audio backbone of the model to a set of music audio classification and regression tasks. We demonstrate the usefulness of our approach by comparing the performance of audio representations produced by the same audio backbone with different training strategies and show that our pre-training method consistently achieves comparable or higher scores on all tasks and datasets considered. Our experiments also confirm that MuLaP effectively leverages audio-caption pairs to learn representations that are competitive with audio-only and cross-modal self-supervised methods in the literature.
Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas
ICASSP4
2022 Violinist Identification Using Note-Level Timbre Feature Distributions
abstract
Modelling musical performers’ individual playing styles based on audio features is important for music education, music expression analysis and music generation. In violin performance, the perception of playing styles are mainly affected by the characteristic musical timbre, which is mostly determined by performers, instruments and recording conditions. To verify if timbre features can describe a performer’s style adequately, we examine a violinist identification method based on note-level timbre feature distributions. We first apply it using solo datasets to recognise professional violinists, then use it to identify master players from commercial concerto recordings. The results show that the designed features and method work very well for both datasets. The identification accuracy with the solo dataset using MFCCs and spectral constrast features are 0.94 and 0.91 respectively. Significantly lower but promising results are reported with the concerto dataset. Results suggest that the selected timbre features can model performers’ individual playing reasonably objectively, regardless of the instrument they play.
Yudong Zhao, György Fazekas, Mark B. Sandler
ICASSP2
2022 The Jazz Ontology: A semantic model and large-scale RDF repositories for jazz
abstract
Jazz is a musical tradition that is just over 100 years old; unlike in other Western musical traditions, improvisation plays a central role in jazz. Modelling the domain of jazz poses some ontological challenges due to specificities in musical content and performance practice, such as band lineup fluidity and importance of short melodic patterns for improvisation. This paper presents the Jazz Ontology – a semantic model that addresses these challenges. Additionally, the model also describes workflows for annotating recordings with melody transcriptions and for pattern search. The Jazz Ontology incorporates existing standards and ontologies such as FRBR and the Music Ontology. The ontology has been assessed by examining how well it supports describing and merging existing datasets and whether it facilitates novel discoveries in a music browsing application. The utility of the ontology is also demonstrated in a novel framework for managing jazz related music information. This involves the population of the Jazz Ontology with the metadata from large scale audio and bibliographic corpora (the Jazz Encyclopedia and the Jazz Discography). The resulting RDF datasets were merged and linked to existing Linked Open Data resources. These datasets are publicly available and are driving an online application that is being used by jazz researchers and music lovers for the systematic study of jazz.
Polina Proutskova, Daniel Wolff, György Fazekas, Klaus Frieler, Frank Höger, Olga Velichkina, Gabriel Solis, Tillman Weyde, Martin Pfleiderer, Hélène C. Crayencour, Geoffroy Peeters, Simon Dixon
J. Web Semant.3
2022 The Smart Musical Instruments Ontology
Luca Turchet, Paolo Bouquet, Andrea Molinari, György Fazekas
J. Web Semant.4
2021 MusCaps: Generating Captions for Music Audio
abstract
Content-based music information retrieval has seen rapid progress with the adoption of deep learning. Current approaches to high-level music description typically make use of classification models, such as in auto-tagging or genre and mood classification. In this work, we propose to address music description via audio captioning, defined as the task of generating a natural language description of music audio content in a human-like manner. To this end, we present the first music audio captioning model, MusCaps, consisting of an encoder-decoder with temporal attention. Our method combines convolutional and recurrent neural network architectures to jointly process audio-text inputs through a multimodal encoder and leverages pretraining on audio data to obtain representations that effectively capture and summarise musical features in the input. Evaluation of the generated captions through automatic metrics shows that our method outperforms a baseline designed for non-music audio captioning. Through an ablation study, we unveil that this performance boost can be mainly attributed to pre-training of the audio encoder, while other design choices - modality fusion, decoding strategy and the use of attention - contribute only marginally. Our model represents a shift away from classification-based music description and combines tasks requiring both auditory and linguistic understanding to bridge the semantic gap in music information retrieval11Code available at https://github.com/ilaria-manco/muscaps.
Ilaria Manco, Emmanouil Benetos, Elio Quinton, György Fazekas
IJCNN4
2021 A Modulation Front-End for Music Audio Tagging
abstract
Convolutional Neural Networks have been extensively explored in the task of automatic music tagging. The problem can be approached by using either engineered time-frequency features or raw audio as input. Modulation filter bank representations that have been actively researched as a basis for timbre perception have the potential to facilitate the extraction of perceptually salient features. We explore end-to-end learned front-ends for audio representation learning, ModNet and SincModNet, that incorporate a temporal modulation processing block. The structure is effectively analogous to a modulation filter bank, where the FIR filter center frequencies are learned in a data-driven manner. The expectation is that a perceptually motivated filter bank can provide a useful representation for identifying music features. Our experimental results provide a fully visualisable and interpretable front-end temporal modulation decomposition of raw audio. We evaluate the performance of our model against the state-of-the-art of music tagging on the MagnaTagATune dataset. We analyse the impact on performance for particular tags when time-frequency bands are subsampled by the modulation filters at a progressively reduced rate. We demonstrate that modulation filtering provides promising results for music tagging and feature representation, without using extensive musical domain knowledge in the design of this frontend.
Cyrus Vahidi, Charalampos Saitis, György Fazekas
IJCNN3
2020 Poster: Programming Practices Among Interactive Audio Software Developers
abstract
New domain-specific languages for creating music and audio applications have typically been created in response to some technological challenge. Recent research has begun looking at how these languages impact our creative and aesthetic choices in music-making but we have little understanding on their effect on our wider programming practice. We present a survey that seeks to uncover what programming practices exist among interactive audio software developers and discover it is highly multi-practice, with developers adopting both exploratory programming and software engineering practice. A Q methodological study reveals that this multi-practice development is supported by different combinations of language features.
Andrew Thompson 0001, György Fazekas, Geraint A. Wiggins
VL/HCC2
2020 The Internet of Audio Things: State of the Art, Vision, and Challenges
abstract
The Internet of Audio Things (IoAuT) is an emerging research field positioned at the intersection of the Internet of Things, sound and music computing, artificial intelligence, and human-computer interaction. The IoAuT refers to the networks of computing devices embedded in physical objects (Audio Things) dedicated to the production, reception, analysis, and understanding of audio in distributed environments. Audio Things, such as nodes of wireless acoustic sensor networks, are connected by an infrastructure that enables multidirectional communication, both locally and remotely. In this article, we first review the state of the art of this field, then we present a vision for the IoAuT and its motivations. In the proposed vision, the IoAuT enables the connection of digital and physical domains by means of appropriate information and communication technologies, fostering novel applications and services based on auditory information. The ecosystems associated with the IoAuT include interoperable devices and services that connect humans and machines to support human-human and human-machines interactions. We discuss the challenges and implications of this field, which lead to future research directions on the topics of privacy, security, design of Audio Things, and methods for the analysis and representation of audio-related information.
Luca Turchet, György Fazekas, Mathieu Lagrange, Hossein Shokri Ghadikolaei, Carlo Fischione
IEEE Internet Things J.2
2020 Cloud-smart Musical Instrument Interactions: Querying a Large Music Collection with a Smart Guitar
abstract
Large online music databases under Creative Commons licenses are rarely recorded by well-known artists, therefore conventional metadata-based search is insufficient in their adaptation to instrument players’ needs. The emerging class of smart musical instruments (SMIs) can address this challenge. Thanks to direct internet connectivity and embedded processing, SMIs can send requests to repositories and reproduce the response for improvisation, composition, or learning purposes. We present a smart guitar prototype that allows retrieving songs from large online music databases using criteria different from conventional music search, which were derived from interviewing 30 guitar players. We investigate three interaction methods coupled with four search criteria (tempo, chords, key and tuning) exploiting intelligent capabilities in the instrument: (i) keywords-based retrieval using an embedded touchscreen; (ii) cloud-computing where recorded content is transmitted to a server that extracts relevant audio features; (iii) edge-computing where the guitar detects audio features and sends the request directly. Overall, the evaluation of these methods with beginner, intermediate, and expert players showed a strong appreciation for the direct connectivity of the instrument with an online database and the approach to the search based on the actual musical content rather than conventional textual criteria, such as song title or artist name.
Luca Turchet, Johan Pauwels, Carlo Fischione, György Fazekas
ACM Trans. Internet Things4
2020 The Internet of Musical Things Ontology
Luca Turchet, Francesco Antoniazzi, Fabio Viola, Fausto Giunchiglia, György Fazekas
J. Web Semant.5
2019 Piano Sustain-pedal Detection Using Convolutional Neural Networks
abstract
Recent research on piano transcription has focused primarily on note events. Very few studies have investigated pedalling techniques, which form an important aspect of expressive piano music performance. In this paper, we propose a novel method for piano sustain-pedal detection based on Convolutional Neural Networks (CNN). Inspired by different acoustic characteristics at the start (pedal onset) versus during the pedalled segment, two binary classifiers are trained separately to learn both temporal dependencies and timbral features using CNN. Their outputs are fused in order to decide whether a portion in a piano recording is played with the sustain pedal. The proposed architecture and our detection system are assessed using a dataset with frame-wise pedal on/off annotations. An average F1 score of 0.74 is obtained for the test set. The method performs better on pieces of Romantic-era composers, who intended to deliver more colours to the piano sound through pedalling techniques.
Beici Liang, György Fazekas, Mark B. Sandler
ICASSP2
2019 Transfer Learning for Piano Sustain-Pedal Detection
abstract
Detecting piano pedalling techniques in polyphonic music remains a challenging task in music information retrieval. While other piano-related tasks, such as pitch estimation and onset detection, have seen improvement through applying deep learning methods, little work has been done to develop deep learning models to detect playing techniques. In this paper, we propose a transfer learning approach for the detection of sustain-pedal techniques, which are commonly used by pianists to enrich the sound. In the source task, a convolutional neural network (CNN) is trained for learning spectral and temporal contexts when the sustain pedal is pressed using a large dataset generated by a physical modelling virtual instrument. The CNN is designed and experimented through exploiting the knowledge of piano acoustics and physics. This can achieve an accuracy score of 0.98 in the validation results. In the target task, the knowledge learned from the synthesised data can be transferred to detect the sustain pedal in acoustic piano recordings. A concatenated feature vector using the activations of the trained convolutional layers is extracted from the recordings and classified into frame-wise pedal press or release. We demonstrate the effectiveness of our method in acoustic piano recordings of Chopin's music. From the cross-validation results, the proposed transfer learning method achieves an average F-measure of 0.89 and an overall performance of 0.84 obtained using the micro-averaged F-measure. These results outperform applying the pre-trained CNN model directly or the model with a fine-tuned last layer.
Beici Liang, György Fazekas, Mark B. Sandler
IJCNN2
2019 A Feature Learning Siamese Model for Intelligent Control of the Dynamic Range Compressor
abstract
In this paper, a siamese DNN model is proposed to learn the characteristics of the audio dynamic range compressor (DRC). This facilitates an intelligent control system that uses audio examples to configure the DRC, a widely used nonlinear audio signal conditioning technique in the areas of music production, speech communication and broadcasting. Several alternative siamese DNN architectures are proposed to learn feature embeddings that can characterise subtle effects due to dynamic range compression. These models are compared with each other as well as handcrafted features proposed in previous work. The evaluation of the relations between the hyperparameters of DNN and DRC parameters are also provided. The best model is able to produce a universal feature embedding that is capable of predicting multiple DRC parameters simultaneously, which is a significant improvement from our previous research. The feature embedding shows better performance than handcrafted audio features when predicting DRC parameters for both mono-instrument audio loops and polyphonic music pieces.
Di Sheng, György Fazekas
IJCNN2
2018 Feature Design Using Audio Decomposition for Intelligent Control of the Dynamic Range Compressor
abstract
This papeper proposes a method of controlling the dynamic range compressor using sound examples. Our earlier work showed the effectiveness of random forest regression to map acoustic features to effect control parameters [1]. We extend this work to address the challenging task of extracting relevant features when audio events overalp. We assess different audio decomposition approaches suchs as onset event detection, NMF, and transient/stationary audio separation using ISTA and compare feature extraction strategies for each case. Numerical and perceptual similarty tests show the utility of audio decomposition as well as specific features in the prediction of dynamic range compressor parameters.
Di Sheng, György Fazekas
ICASSP2
2018 Audio Commons Ontology: A Data Model for an Audio Content Ecosystem
Miguel Ceriani, György Fazekas
ISWC (2)2
2017 Convolutional recurrent neural networks for music classification
abstract
We introduce a convolutional recurrent neural network (CRNN) for music tagging. CRNNs take advantage of convolutional neural networks (CNNs) for local feature extraction and recurrent neural networks for temporal summarisation of the extracted features. We compare CRNN with three CNN structures that have been used for music tagging while controlling the number of parameters with respect to their performance and training time per sample. Overall, we found that CRNNs show a strong performance with respect to the number of parameter and training time, indicating the effectiveness of its hybrid structure in music feature extraction and feature summarisation.
Keunwoo Choi, György Fazekas, Mark B. Sandler, Kyunghyun Cho
ICASSP2
2017 Linked Data Publication of Live Music Archives and Analyses
Sean Bechhofer, Kevin R. Page, David M. Weigl, György Fazekas, Thomas Wilmering
ISWC (2)4
2016 Hybrid music recommender using content-based and social information
abstract
Internet resources available today, including songs, albums, playlists or podcasts, that a user cannot discover if there is not a tool to filter the items that the user might consider relevant. Several recommendation techniques have been developed since the Internet explosion to achieve this filtering task. In an attempt to recommend relevant songs to users, we propose an hybrid recommender that considers real-world users information and high-level representation for audio data. We use a deep learning technique, convolutional deep neural networks, to represent an audio segment in a n-dimensional vector, whose dimensions define the probability of the segment to belong to a specific music genre. To capture the listening behavior of a user, we investigate a state-of-the-art technique, estimation of distribution algorithms. The designed hybrid music recommender outperforms the predictions compared with a traditional content-based recommender.
Paulo Chiliguano, György Fazekas
ICASSP2
2016 Semantic Description of Timbral Transformations in Music Production
abstract
In music production, descriptive terminology is used to define perceived sound transformations. By understanding the underlying statistical features associated with these descriptions, we can aid the retrieval of contextually relevant processing parameters using natural language, and create intelligent systems capable of assisting in audio engineering. In this study, we present an analysis of a dataset containing descriptive terms gathered using a series of processing modules, embedded within a Digital Audio Workstation. By applying hierarchical clustering to the audio feature space, we show that similarity in term representations exists within and between transformation classes. Furthermore, the organisation of terms in low-dimensional timbre space can be explained using perceptual concepts such as size and dissonance. We conclude by performing Latent Semantic Indexing to show that similar groupings exist based on term frequency.
Ryan Stables, Brecht De Man, Sean Enderby, Joshua D. Reiss, György Fazekas, Thomas Wilmering
ACM Multimedia5
2016 Ontological Representation of Audio Features
Alo Allik, György Fazekas, Mark B. Sandler
ISWC (2)2
2016 AUFX-O: Novel Methods for the Representation of Audio Processing Workflows
Thomas Wilmering, György Fazekas, Mark B. Sandler
ISWC (2)2
2016 Genre-Adaptive Semantic Computing and Audio-Based Modelling for Music Mood Annotation
abstract
This study investigates whether taking genre into account is beneficial for automatic music mood annotation in terms of core affects valence, arousal, and tension, as well as several other mood scales. Novel techniques employing genre-adaptive semantic computing and audio-based modelling are proposed. A technique called the ACTwg employs genre-adaptive semantic computing of mood-related social tags, whereas ACTwg-SLPwg combines semantic computing and audio-based modelling, both in a genre-adaptive manner. The proposed techniques are experimentally evaluated at predicting listener ratings related to a set of 600 popular music tracks spanning multiple genres. The results show that ACTwg outperforms a semantic computing technique that does not exploit genre information, and ACTwg-SLPwg outperforms conventional techniques and other genre-adaptive alternatives. In particular, improvements in the prediction rates are obtained for the valence dimension which is typically the most challenging core affect dimension for audio-based annotation. The specificity of genre categories is not crucial for the performance of ACTwg-SLPwg. The study also presents analytical insights into inferring a concise tag-based genre representation for genre-adaptive music mood analysis.
Pasi Saari, György Fazekas, Tuomas Eerola, Mathieu Barthet, Olivier Lartillot, Mark B. Sandler
IEEE Trans. Affect. Comput.2
2015 On the use of the tempogram to describe audio content and its application to Music structural segmentation
abstract
This paper presents a new set of audio features to describe music content based on tempo cues. Tempogram, a mid-level representation of tempo information, is constructed to characterize tempo variation and local pulse in the audio signal. We introduce a collection of novel tempogram-based features inspired by musicological hypotheses about the relation between music structure and its rhythmic components prominent at different metrical levels. The strength of these features is demonstrated in music structural segmentation, an important task in Music information retrieval (MIR), using several published popular music datasets. Results indicate that incorporating tempo information into audio segmentation is a promising new direction.
Mi Tian 0001, György Fazekas, Dawn A. A. Black, Mark B. Sandler
ICASSP2
2013 Mood Conductor: Emotion-Driven Interactive Music Performance
abstract
Mood Conductor is a system that allows the audience to interact with stage performers to create directed improvisations. The term "conductor" is used metaphorically. Rather than directing a musical performance by way of visible gestures, spectators act as conductors by communicating emotional intentions to the performers through our web-based smartphone-friendly Mood Conductor app. Performers receive the audience's directions via a visual feedback system operating in real-time. Emotions are represented by coloured blobs in a two-dimensional space (vertical dimension: arousal or excitation; horizontal dimension: valence or pleasantness). The size of the "emotion blobs" indicates the number of spectators that have selected the corresponding emotions at a given time.
György Fazekas, Mathieu Barthet, Mark B. Sandler
ACII1
2013 Towards the Representation of Chinese Traditional Music: A State of the Art Review of Music Metadata Standards
Mi Tian 0001, György Fazekas, Dawn A. A. Black, Mark B. Sandler
Dublin Core Conference2
2013 Automatic Ontology Generation for Musical Instruments Based on Audio Analysis
abstract
In this paper we present a novel hybrid system that involves a formal method of automatic ontology generation for web-based audio signal processing applications. An ontology is seen as a knowledge management structure that represents domain knowledge in a machine interpretable format. It describes concepts and relationships within a particular domain, in our case, the domain of musical instruments. However, the different tasks of ontology engineering including manual annotation, hierarchical structuring and organization of data can be laborious and challenging. For these reasons, we investigate how the process of creating ontologies can be made less dependent on human supervision by exploring concept analysis techniques in a Semantic Web environment. In this study, various musical instruments, from wind to string families, are classified using timbre features extracted from audio. To obtain models of the analysed instrument recordings, we use K-means clustering to determine an optimised codebook of Line Spectral Frequencies (LSFs), or Mel-frequency Cepstral Coefficients (MFCCs). Two classification techniques based on Multi-Layer Perceptron (MLP) neural network and Support Vector Machines (SVM) were tested. Then, Formal Concept Analysis (FCA) is used to automatically build the hierarchical structure of musical instrument ontologies. Finally, the generated ontologies are expressed using the Ontology Web Language (OWL). System performance was evaluated under natural recording conditions using databases of isolated notes and melodic phrases. Analysis of Variance (ANOVA) were conducted with the feature and classifier attributes as independent variables and the musical instrument recognition F-measure as dependent variable. Based on these statistical analyses, a detailed comparison between musical instrument recognition models is made to investigate their effects on the automatic ontology generation system. The proposed system is general and also applicable to other research fields that are related to ontologies and the Semantic Web.
Sefki Kolozali, Mathieu Barthet, György Fazekas, Mark B. Sandler
IEEE Trans. Speech Audio Process.3