Matt McVicar

dblp:96/9888 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
1since 2021 · last 2024
0000-0003-0212-4093ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 4 first-authorGraphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Audio and music processing · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › music transcription
chord recognition
0.532014
Automatic Chord Estimation from Audio: A Review of the State of the Art · IEEE ACM Trans. Audio Speech Lang. Process. 2014
Understanding Effects of Subjectivity in Measuring Chord Estimation Accuracy · IEEE ACM Trans. Audio Speech Lang. Process. 2013
An End-to-End Machine Learning System for Harmonic Analysis of Music · IEEE Trans. Speech Audio Process. 2012
Audio and music processing
music analysis
0.322013
Understanding Effects of Subjectivity in Measuring Chord Estimation Accuracy · IEEE ACM Trans. Audio Speech Lang. Process. 2013
An End-to-End Machine Learning System for Harmonic Analysis of Music · IEEE Trans. Speech Audio Process. 2012
Audio and music processing
music generation
0.212015
AutoGuitarTab: Computer-Aided Composition of Rhythm and Lead Guitar Parts in the Tablature Space · IEEE ACM Trans. Audio Speech Lang. Process. 2015
Audio and music processing
music information retrieval
0.212014
Automatic Chord Estimation from Audio: A Review of the State of the Art · IEEE ACM Trans. Audio Speech Lang. Process. 2014

Methods — techniques the papers use, named apart from their topics

user survey · 0.2structural analysis · 0.2markov chain · 0.2annotation subjectivity analysis · 0.2machine learning · 0.1chromagram · 0.1
YearPublicationVenuePosition
2024 Resource-Constrained Stereo Singing Voice Cancellation
abstract
We study the problem of stereo singing voice cancellation, a subtask of music source separation, whose goal is to estimate an instrumental background from a stereo mix. We explore how to achieve performance similar to large state-of-the-art source separation networks starting from a small, efficient model for real-time speech separation. Such a model is useful when memory and compute are limited and singing voice processing has to run with limited look-ahead. In practice, this is realised by adapting an existing mono model to handle stereo input. Improvements in quality are obtained by tuning model parameters and expanding the training set. Moreover, we highlight the benefits a stereo model brings by introducing a new metric which detects attenuation inconsistencies between channels. Our approach is evaluated using objective offline metrics and a large-scale MUSHRA trial, confirming the effectiveness of our techniques in stringent listening tests.
Clara Borrelli, James Rae, Dogac Basaran, Matt McVicar, Mehrez Souden, Matthias Mauch
ICASSP4
2017 Hierarchical Novelty Detection
Paolo Simeone, Raúl Santos-Rodríguez, Matt McVicar, Jefrey Lijffijt, Tijl De Bie
IDA3
2016 Learning to separate vocals from polyphonic mixtures via ensemble methods and structured output prediction
abstract
Separating the singing from a polyphonic mixed audio signal is a challenging but important task, with a wide range of applications across the music industry and music informatics research. Various methods have been devised over the years, ranging from Deep Learning approaches to dedicated ad hoc solutions. In this paper, we present a novel machine learning method for the task, using a Conditional Random Field (CRF) approach for structured output prediction. We exploit the diversity of previously proposed approaches by using their predictions as input features to our method - thus effectively developing an ensemble method. Our empirical results demonstrate the potential of integrating predictions from different previously-proposed methods into one ensemble method, and additionally show that CRF models with larger complexities generally lead to superior performance.
Matt McVicar, Raúl Santos-Rodríguez, Tijl De Bie
ICASSP1
2016 SuMoTED: An intuitive edit distance between rooted unordered uniquely-labelled trees
abstract
Defining and computing distances between tree structures is a classical area of study in theoretical computer science, with practical applications in the areas of computational biology, information retrieval, text analysis, and many others. In this paper, we focus on rooted, unordered, uniquely-labelled trees such as taxonomies and other hierarchies. For trees as these, we introduce the intuitive concept of a ‘local move’ operation as an atomic edit of a tree. We then introduce SuMoTED, a new edit distance measure between such trees, defined as the minimal number of local moves required to convert one tree into another. We show how SuMoTED can be computed using a scalable algorithm with quadratic time complexity. Finally, we demonstrate its use on a collection of music genre taxonomies.
Matt McVicar, Benjamin Sach, Cédric Mesnage, Jefrey Lijffijt, Eirini Spyropoulou, Tijl De Bie
Pattern Recognit. Lett.1
2015 Interactively Exploring Supply and Demand in the UK Independent Music Scene
Matt McVicar, Cédric Mesnage, Jefrey Lijffijt, Tijl De Bie
ECML/PKDD (3)1
2015 AutoGuitarTab: Computer-Aided Composition of Rhythm and Lead Guitar Parts in the Tablature Space
abstract
We present AutoGuitarTab, a system for generating realistic guitar tablature given an input symbolic chord and key sequence. Our system consists of two modules: AutoRhythmGuitar and AutoLeadGuitar. The first of these generates rhythm guitar tablatures which outline the input chord sequence in a particular style (using Markov chains to ensure playability) and performs a structural analysis to produce a structurally consistent composition. AutoLeadGuitar generates lead guitar parts in distinct musical phrases, guiding the pitch classes towards chord tones and steering the evolution of the rhythmic and melodic intensity according to user preference. Experimentally, we uncover musician-specific trends in guitar playing style, and demonstrate our system's ability to produce playable, realistic and style-specific tablature using a combination of algorithmic, user-surveyed and expert evaluation techniques.
Matt McVicar, Satoru Fukayama, Masataka Goto
IEEE ACM Trans. Audio Speech Lang. Process.1
2014 Leveraging repetition for improved automatic lyric transcription in popular music
abstract
Transcribing lyrics from musical audio is a challenging research problem which has not benefited from many advances made in the related field of automatic speech recognition, owing to the prevalent musical accompaniment and differences between the spoken and sung voice. However, one aspect of this problem which has yet to be exploited by researchers is that significant portions of the lyrics will be repeated throughout the song. In this paper we investigate how this information can be leveraged to form a consensus transcription with improved consistency and accuracy. Our results show that improvements can be gained using a variety of techniques, and that relative gains are largest under the most challenging and realistic experimental conditions.
Matt McVicar, Daniel P. W. Ellis, Masataka Goto
ICASSP1
2014 Automatic Chord Estimation from Audio: A Review of the State of the Art
abstract
In this overview article, we review research on the task of Automatic Chord Estimation (ACE). The major contributions from the last 14 years of research are summarized, with detailed discussions of the following topics: feature extraction, modeling strategies, model training and datasets, and evaluation strategies. Results from the annual benchmarking evaluation Music Information Retrieval Evaluation eXchange (MIREX) are also discussed as well as developments in software implementations and the impact of ACE within MIR. We conclude with possible directions for future research.
Matt McVicar, Raúl Santos-Rodríguez, Yizhao Ni, Tijl De Bie
IEEE ACM Trans. Audio Speech Lang. Process.1
2013 Understanding Effects of Subjectivity in Measuring Chord Estimation Accuracy
abstract
To assess the performance of an automatic chord estimation system, reference annotations are indispensable. However, owing to the complexity of music and the sometimes ambiguous harmonic structure of polyphonic music, chord annotations are inherently subjective, and as a result any derived accuracy estimates will be subjective as well. In this paper, we investigate the extent of the confounding effect of subjectivity in reference annotations. Our results show that this effect is important, and they affect different types of automatic chord estimation systems in different ways. Our results have implications for research on automatic chord estimation, but also on other fields that evaluate performance by comparing against human provided annotations that are confounded by subjectivity.
Yizhao Ni, Matt McVicar, Raúl Santos-Rodríguez, Tijl De Bie
IEEE ACM Trans. Audio Speech Lang. Process.2
2012 An End-to-End Machine Learning System for Harmonic Analysis of Music
abstract
We present a new system for the harmonic analysis of popular musical audio. It is focused on chord estimation, although the proposed system additionally estimates the key sequence and bass notes. It is distinct from competing approaches in two main ways. First, it makes use of a new improved chromagram representation of audio that takes the human perception of loudness into account. Furthermore, it is the first system for joint estimation of chords, keys, and bass notes that is fully based on machine learning, requiring no expert knowledge to tune the parameters. This means that it will benefit from future increases in available annotated audio files, broadening its applicability to a wider range of genres. In all of three evaluation scenarios, including a new one that allows evaluation on audio for which no complete ground truth annotation is available, the proposed system is shown to be faster, more memory efficient, and more accurate than the state-of-the-art.
Yizhao Ni, Matt McVicar, Raúl Santos-Rodríguez, Tijl De Bie
IEEE Trans. Speech Audio Process.2