Juan Pablo Bello

dblp:56/4223 · DBLP profile ↗
← Back
61ranked-venue papers
6as first author
22since 2021 · last 2025
0000-0001-8561-5204ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 45 · 3 first-author · 19 since 2021Artificial intelligence and machine learning · 13 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 3 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Towards Few-Shot Training-Free Anomaly Sound Detection
Ho-Hsiang Wu, Abinaya Kumar, Luca Bondi, Shabnam Ghaffarzadegan, Juan Pablo Bello
INTERSPEECH6
2025 Perceptually-Guided Acoustic "Foveation"
abstract
Realistic spatial audio rendering improves immersion in virtual environments. However, the computational complexity of acoustic propagation increases linearly with the number of sources. Consequently, real-time accurate acoustic rendering becomes challenging in highly dynamic scenarios such as virtual and augmented reality (VR/AR). Exploiting the fact that human spatial sensitivity of acoustic sources is not equal at azimuth eccentricities in the horizontal plane, we introduce a perceptually-aware acoustic "foveation" guidance model to the audio rendering pipeline, which can integrate audio sources that are not spatially resolvable by human listeners. To this end, we first conduct a series of psychophysical studies to measure the minimum resolvable audible angular distance under various spatial and background conditions. We leverage this data to derive an azimuth-characterized real-time acoustic foveation algorithm. Numerical analysis and subjective user studies in VR environments demonstrate our method’s effectiveness in significantly reducing acoustic rendering workload, without compromising users’ spatial perception of audio sources. We believe that the presented research will motivate future investigation into the new frontier of modeling and leveraging human multimodal perceptual limitations — beyond the extensively studied visual acuity — for designing efficient VR/AR systems.
Kenneth Chen, Irán R. Román, Juan Pablo Bello, Qi Sun 0003, Praneeth Chakravarthula
VR4
2025 HuBar: A Visual Analytics Tool to Explore Human Behavior Based on fNIRS in AR Guidance Systems
abstract
The concept of an intelligent augmented reality (AR) assistant has significant, wide-ranging applications, with potential uses in medicine, military, and mechanics domains. Such an assistant must be able to perceive the environment and actions, reason about the environment state in relation to a given task, and seamlessly interact with the task performer. These interactions typically involve an AR headset equipped with sensors which capture video, audio, and haptic feedback. Previous works have sought to facilitate the development of intelligent AR assistants by visualizing these sensor data streams in conjunction with the assistant's perception and reasoning model outputs. However, existing visual analytics systems do not focus on user modeling or include biometric data, and are only capable of visualizing a single task session for a single performer at a time. Moreover, they typically assume a task involves linear progression from one step to the next. We propose a visual analytics system that allows users to compare performance during multiple task sessions, focusing on non-linear tasks where different step sequences can lead to success. In particular, we design visualizations for understanding user behavior through functional near-infrared spectroscopy (fNIRS) data as a proxy for perception, attention, and memory as well as corresponding motion data (acceleration, angular velocity, and gaze). We distill these insights into embedding representations that allow users to easily select groups of sessions with similar behaviors. We provide two case studies that demonstrate how to use these visualizations to gain insights about task performance using data collected during helicopter copilot training tasks. Finally, we evaluate our approach through an in-depth examination of a think-aloud experiment with five domain experts.
Sonia Castelo Quispe, João Rulff, Parikshit Solunke, Erin McGowan, Guande Wu, Irán R. Román, Roque Lopez, Bea Steers, Qi Sun 0003, Juan Pablo Bello, Bradley Feest, Michael Middleton, Ryan McKendrick, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.10
2024 Spatial Scaper: A Library to Simulate and Augment Soundscapes for Sound Event Localization and Detection in Realistic Rooms
abstract
Sound event localization and detection (SELD) is an important task in machine listening. Major advancements rely on simulated data with sound events in specific rooms and strong spatio-temporal labels. SELD data is simulated by convolving spatialy-localized room impulse responses (RIRs) with sound waveforms to place sound events in a soundscape. However, RIRs require manual collection in specific rooms. We present SpatialScaper, a library for SELD data simulation and augmentation. Compared to existing tools, SpatialScaper emulates virtual rooms via parameters such as size and wall absorption. This allows for parameterized placement (including movement) of foreground and background sound sources. SpatialScaper also includes data augmentation pipelines that can be applied to existing SELD data. As a case study, we use SpatialScaper to add rooms to the DCASE SELD data. Training a model with our data led to progressive performance improves as a direct function of acoustic diversity. These results show that SpatialScaper is valuable to train robust SELD models.
Irán R. Román, Christopher Ick, Sivan Ding, Adrian S. Roman, Brian McFee, Juan Pablo Bello
ICASSP6
2024 Robust DoA Estimation from Deep Acoustic Imaging
abstract
Direction of arrival estimation (DoAE) aims at tracking a sound in azimuth and elevation. Recent advancements include data-driven models with inputs derived from ambisonics intensity vectors or correlations between channels in a microphone array. A spherical intensity map (SIM), or acoustic image, is an alternative input representation that remains underexplored. SIMs benefit from high-resolution microphone arrays, yet most DoAE datasets use low-resolution ones. Therefore, we first propose a super-resolution method to upsample low-resolution microphones. Next, we benchmark DoAE models that use SIMs as input. We arrive to a model that uses SIMs for DoAE estimation and outperforms a baseline and a state-of-the-art model. Our study highlights the relevance of acoustic imaging for DoAE tasks.
Adrian S. Roman, Irán R. Román, Juan Pablo Bello
ICASSP3
2024 BirdVoxDetect: Large-Scale Detection and Classification of Flight Calls for Bird Migration Monitoring
abstract
Sound event classification has the potential to advance our understanding of bird migration. Although it is long known that migratory species have a vocal signature of their own, previous work on automatic flight call classification has been limited in robustness and scope: e.g., covering few recording sites, short acquisition segments, and simplified biological taxonomies. In this paper, we present BirdVoxDetect (BVD), the first full-fledged solution to bird migration monitoring from acoustic sensor network data. As an open-source software, BVD integrates an original pipeline of three machine learning modules. The first module is a random forest classifier of sensor faults, trained with human-in-the-loop active learning. The second module is a deep convolutional neural network for sound event detection with per-channel energy normalization (PCEN). The third module is a multitask convolutional neural network which predicts the family, genus, and species of flight calls from passerines(Passeriformes)of North America. We evaluate BVD on a new dataset (296 hours from nine locations, the largest to date for this task) and discuss the main sources of estimation error in a real-world deployment: mechanical sensor failures, sensitivity to background noise, misdetection, and taxonomic confusion. Then, we deploy BVD to an unprecedented scale: 6672 hours of audio (approximately one terabyte), corresponding to a full season of bird migration. Running BVD in parallel over the full-season dataset yields 1.6 billion FFT's, 480 million neural network predictions, and over six petabytes of throughput. With this method, our main finding is that deep learning and bioacoustic sensor networks are ready to complement radar observations and crowdsourced surveys for bird migration monitoring, thus benefiting conservation ecology and land-use planning at large.
Vincent Lostanlen, Aurora Cramer, Justin Salamon, Andrew Farnsworth, Benjamin Van Doren, Steve Kelling, Juan Pablo Bello
IEEE ACM Trans. Audio Speech Lang. Process.7
2024 : Visualization of AI-Assisted Task Guidance in AR
abstract
The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant.
Sonia Castelo Quispe, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao 0013, Kyunghyun Cho, He He 0001, Qi Sun 0003, Huy T. Vo, Juan Pablo Bello, Michael Krone, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.16
2023 Does a Quieter City Mean Fewer Complaints? The Sounds of New York City During Covid-19 Lockdown
abstract
The COVID-19 pandemic had an unprecedented effect in human activity and city landscapes. A very notorious transformation during this period was the change in noise levels and patterns across cities. Small scale studies have show this change in noise levels across different locations in the globe. In this work, we extend these studies by using historical audio data from the SONYC sensor network deployed in New York City. We exploit machine listening models to understand not only noise levels but also patterns, by performing a sound source presence analysis. Finally, we contrast our finding from the acoustic data with noise complaints to better understand the relationship between noise and our perception of it.
Mark Cartwright, Magdalena Fuentes, Charlie Mydlarz, Fabio Miranda 0001, Juan Pablo Bello
ICASSP5
2023 Exploring Approaches to Multi-Task Automatic Synthesizer Programming
abstract
Automatic Synthesizer Programming is the task of transforming an audio signal that was generated from a virtual instrument, into the parameters of a sound synthesizer that would generate this signal. In the past, this could only be done for one virtual instrument. In this paper, we expand the current literature by exploring approaches to automatic synthesizer programming for multiple virtual instruments. Two different approaches to multi-task automatic synthesizer programming are presented. We find that the joint-decoder approach performs best. We also evaluate the performance of this model for different timbre instruments and different latent dimension sizes.
Daniel Faronbi, Irán R. Román, Juan Pablo Bello
ICASSP3
2023 Flowgrad: Using Motion for Visual Sound Source Localization
abstract
Most recent work in visual sound source localization relies on semantic audio-visual representations learned in a self-supervised manner and, by design, excludes temporal information present in videos. While it proves to be effective for widely used benchmark datasets, the method falls short for challenging scenarios like urban traffic. This work introduces temporal context into the state-of-the-art methods for sound source localization in urban scenes using optical flow to encode motion information. An analysis of the strengths and weaknesses of our methods helps us better understand the problem of visual sound source localization and sheds light on open challenges for audio-visual scene understanding. The code and pretrained models are publicly available at https://github.com/rrrajjjj/flowgrad
Rajsuryan Singh, Pablo Zinemanas, Xavier Serra, Juan Pablo Bello, Magdalena Fuentes
ICASSP4
2023 Audio-Text Models Do Not Yet Leverage Natural Language
abstract
Multi-modal contrastive learning techniques in the audio-text domain have quickly become a highly active area of research. Most works are evaluated with standard audio retrieval and classification benchmarks assuming that (i) these models are capable of leveraging the rich information contained in natural language, and (ii) current benchmarks are able to capture the nuances of such information. In this work, we show that state-of-the-art audio-text models do not yet really understand natural language, especially contextual concepts such as sequential or concurrent ordering of sound events. Our results suggest that existing benchmarks are not sufficient to assess these models' capabilities to match complex contexts from the audio and text modalities. We propose a Transformer-based architecture and show that, unlike prior work, it is capable of modeling the sequential relationship between sound events in the text and audio, given appropriate benchmark data. We advocate for the collection or generation of additional, diverse, data to allow future research to fully leverage natural language for audio-text modeling.
Ho-Hsiang Wu, Oriol Nieto, Juan Pablo Bello, Justin Salamon
ICASSP3
2022 Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
abstract
Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To address this we present a curated audio-visual dataset, Urban Sound & Sight (Urbansas), developed for investigating the detection and localization of sounding vehicles in the wild. Urbansas consists of 12 hours of unlabeled data along with 3 hours of manually annotated data, including bounding boxes with classes and unique id of vehicles, and strong audio labels featuring vehicle types and indicating off-screen sounds. We discuss the challenges presented by the dataset and how to use its annotations for the localization of vehicles in the wild through audio models.
Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, Juan Pablo Bello
ICASSP11
2022 Few-Shot Musical Source Separation
abstract
Deep learning-based approaches to musical source separation are often limited to the instrument classes that the models are trained on and do not generalize to separate unseen instruments. To address this, we propose a few-shot musical source separation paradigm. We condition a generic U-Net source separation model using few audio examples of the target instrument. We train a few-shot conditioning encoder jointly with the U-Net to encode the audio examples into a conditioning vector to configure the U-Net via feature-wise linear modulation (FiLM). We evaluate the trained models on real musical recordings in the MUSDB18 and MedleyDB datasets. We show that our proposed few-shot conditioning paradigm outperforms the base-line one-hot instrument-class conditioned model for both seen and unseen instruments. To extend the scope of our approach to a wider variety of real-world scenarios, we also experiment with different conditioning example characteristics, including examples from different recordings, with multiple sources, or negative conditioning examples.
Yu Wang 0105, Daniel Stoller, Rachel M. Bittner, Juan Pablo Bello
ICASSP4
2022 Wav2CLIP: Learning Robust Audio Representations from Clip
abstract
We propose Wav2CLIP, a robust audio representation learning method by distilling from Contrastive Language-Image Pre-training (CLIP). We systematically evaluate Wav2CLIP on a variety of audio tasks including classification, retrieval, and generation, and show that Wav2CLIP can outperform several publicly available pre-trained audio representation algorithms. Wav2CLIP projects audio into a shared embedding space with images and text, which enables multimodal applications such as zero-shot classification, and cross-modal retrieval. Furthermore, Wav2CLIP needs just ∼10% of the data to achieve competitive performance on downstream tasks compared with fully supervised models, and is more efficient to pre-train than competing methods as it does not require learning a visual model in concert with an auditory model. Finally, we demonstrate image generation from Wav2CLIP as qualitative assessment of the shared embedding space. Our code and model weights are open sourced and made available for further applications.
Ho-Hsiang Wu, Prem Seetharaman, Juan Pablo Bello
ICASSP4
2022 Active Few-Shot Learning for Sound Event Detection
Yu Wang 0105, Mark Cartwright, Juan Pablo Bello
INTERSPEECH3
2022 How to Listen? Rethinking Visual Sound Localization
abstract
Localizing visual sounds consists on locating the position of objects that emit sound within an image.It is a growing research area with potential applications in monitoring natural and urban environments, such as wildlife migration and urban traffic.Previous works are usually evaluated with datasets having mostly a single dominant visible object, and proposed models usually require the introduction of localization modules during training or dedicated sampling strategies, but it remains unclear how these design choices play a role in the adaptability of these methods in more challenging scenarios.In this work, we analyze various model choices for visual sound localization and discuss how their different components affect the model's performance, namely the encoders' architecture, the loss function and the localization strategy.Furthermore, we study the interaction between these decisions, the model performance, and the data, by digging into different evaluation datasets spanning different difficulties and characteristics, and discuss the implications of such decisions in the context of real-world applications.Our code and model weights are open-sourced and made available for further applications.
Ho-Hsiang Wu, Magdalena Fuentes, Prem Seetharaman, Juan Pablo Bello
INTERSPEECH4
2022 Urban Rhapsody: Large-scale exploration of urban soundscapes
abstract
Abstract Noise is one of the primary quality‐of‐life issues in urban environments. In addition to annoyance, noise negatively impacts public health and educational performance. While low‐cost sensors can be deployed to monitor ambient noise levels at high temporal resolutions, the amount of data they produce and the complexity of these data pose significant analytical challenges. One way to address these challenges is through machine listening techniques, which are used to extract features in attempts to classify the source of noise and understand temporal patterns of a city's noise situation. However, the overwhelming number of noise sources in the urban environment and the scarcity of labeled data makes it nearly impossible to create classification models with large enough vocabularies that capture the true dynamism of urban soundscapes. In this paper, we first identify a set of requirements in the yet unexplored domain of urban soundscape exploration. To satisfy the requirements and tackle the identified challenges, we propose Urban Rhapsody, a framework that combines state‐of‐the‐art audio representation, machine learning and visual analytics to allow users to interactively create classification models, understand noise patterns of a city, and quickly retrieve and label audio excerpts in order to create a large high‐precision annotated database of urban sound recordings. We demonstrate the tool's utility through case studies performed by domain experts using data generated over the five‐year deployment of a one‐of‐a‐kind sensor network in New York City.
João Rulff, Fabio Miranda 0001, Marcos Lage, Mark Cartwright, Graham Dove, Juan Pablo Bello, Cláudio T. Silva
Comput. Graph. Forum7
2022 From Environmental Monitoring to Mitigation Action: Considerations, Challenges, and Opportunities for HCI
abstract
Noise pollution is among the most consistently cited and highest impact quality-of-life issues in major urban areas across the US, with more than 70 million people estimated to be exposed to noise levels considered harmful. While HCI and CSCW has a relatively rich history of engagement with monitoring such environmental concerns, e.g. through participatory sensing, prior research has not to our knowledge engaged with the process of municipal mitigation. In this paper we present research in this direction, connecting support for pervasive environmental monitoring with civic engagement in mitigation action. Having first identified and described the research space for this engagement, we present empirical data from an ongoing case study focused on two communities living with chronic problem noise. We probe the experiences of residents and representatives of different municipal organizations tasked with addressing their concerns. We find that making and drawing on records of noise is important to residents and authorities, that these groups have misaligned perceptions of how effective current reporting programs are, and that communication between them can be poor. We also find that noise is often one part of more complex issues. We then discuss our findings in light of prior research, and identify a model of civic sensing that highlights opportunities for HCI design and research including: supporting residents' coordinated actions and actions with municipal open data, mediating residents' and authorities' assessments of data quality, and supporting accountability and attributable mitigation action.
Graham Dove, Daniel Fries, Vanessa Johnson, Charlie Mydlarz, Juan Pablo Bello, Oded Nov
Proc. ACM Hum. Comput. Interact.6
2022 Eliciting Confidence for Improving Crowdsourced Audio Annotations
abstract
In this work we explore confidence elicitation methods for crowdsourcing "soft" labels, e.g., probability estimates, to reduce the annotation costs for domains with ambiguous data. Machine learning research has shown that such "soft" labels are more informative and can reduce the data requirements when training supervised machine learning models. By reducing the number of required labels, we can reduce the costs of slow annotation processes such as audio annotation. In our experiments we evaluated three confidence elicitation methods: 1) "No Confidence" elicitation, 2) "Simple Confidence" elicitation, and 3) "Betting" mechanism for confidence elicitation, at both individual (i.e., per participant) and aggregate (i.e., crowd) levels. In addition, we evaluated the interaction between confidence elicitation methods, annotation types (binary, probability, and z-score derived probability), and "soft" versus "hard" (i.e., binarized) aggregate labels. Our results show that both confidence elicitation mechanisms result in higher annotation quality than the "No Confidence" mechanism for binary annotations at both participant and recording levels. In addition, when aggregating labels at the recording level, results indicate that we can achieve comparable results to those with 10-participant aggregate annotations using fewer annotators if we aggregate "soft" labels instead of "hard" labels. These results suggest that for binary audio annotation using a confidence elicitation mechanism and aggregating continuous labels we can obtain higher annotation quality, more informative labels, with quality differences more pronounced with fewer participants. Finally, we propose a way of integrating these confidence elicitation methods into a two-stage, multi-label annotation pipeline.
Ana Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello, Oded Nov
Proc. ACM Hum. Comput. Interact.3
2021 Few-Shot Continual Learning for Audio Classification
abstract
Supervised learning for audio classification typically imposes a fixed class vocabulary, which can be limiting for real-world applications where the target class vocabulary is not known a priori or changes dynamically. In this work, we introduce a few-shot continual learning framework for audio classification, where we can continuously expand a trained base classifier to recognize novel classes based on only few labeled data at inference time. This enables fast and interactive model updates by end-users with minimal human effort. To do so, we leverage the dynamic few-shot learning technique and adapt it to a challenging multi-label audio classification scenario. We incorporate a recent state-of-the-art audio feature extraction model as a backbone and perform a comparative analysis of our approach on two popular audio datasets (ESC-50 and AudioSet). We conduct an in-depth evaluation to illustrate the complexities of the problem and show that, while there is still room for improvement, our method outperforms three baselines on novel class detection while maintaining its performance on base classes.
Yu Wang 0105, Nicholas J. Bryan, Mark Cartwright, Juan Pablo Bello, Justin Salamon
ICASSP4
2021 Specialized Embedding Approximation for Edge Intelligence: A Case Study in Urban Sound Classification
abstract
Embedding models that encode semantic information into low-dimensional vector representations are useful in various machine learning tasks with limited training data. However, these models are typically too large to support inference in small edge devices, which motivates training of smaller yet comparably predictive student embedding models through knowledge distillation (KD). While knowledge distillation traditionally uses the teacher’s original training dataset to train the student, we hypothesize that using a dataset similar to the student’s target domain allows for better compression and training efficiency for the said domain, at the cost of reduced generality across other (non-pertinent) domains. Hence, we introduce Specialized Embedding Approximation (SEA) to train a student featurizer to approximate the teacher’s embedding manifold for a given target domain. We demonstrate the feasibility of SEA in the context of acoustic event classification for urban noise monitoring and show that leveraging a dataset related to this target domain not only improves the baseline performance of the original embedding model but also yields competitive students with >1 order of magnitude lesser storage and activation memory. We further investigate the impact of using random and informed sampling techniques for dimensionality reduction in SEA.
Sangeeta Srivastava, Dhrubojyoti Roy, Mark Cartwright, Juan Pablo Bello, Anish Arora
ICASSP4
2021 Multi-Task Self-Supervised Pre-Training for Music Classification
abstract
Deep learning is very data hungry, and supervised learning especially requires massive labeled data to work well. Machine listening research often suffers from limited labeled data problem, as human annotations are costly to acquire, and annotations for audio are time consuming and less intuitive. Besides, models learned from labeled dataset often embed biases specific to that particular dataset. Therefore, unsupervised learning techniques become popular approaches in solving machine listening problems. Particularly, a self-supervised learning technique utilizing reconstructions of multiple hand-crafted audio features has shown promising results when it is applied to speech domain such as emotion recognition and automatic speech recognition (ASR). In this paper, we apply self-supervised and multi-task learning methods for pre-training music encoders, and explore various design choices including encoder architectures, weighting mechanisms to combine losses from multiple tasks, and worker selections of pretext tasks. We investigate how these design choices interact with various downstream music classification tasks. We find that using various music specific workers altogether with weighting mechanisms to balance the losses during pre-training helps improve and generalize to the downstream tasks.
Ho-Hsiang Wu, Chieh-Chi Kao, Qingming Tang, Ming Sun 0007, Brian McFee, Juan Pablo Bello, Chao Wang 0018
ICASSP6
2020 Chirping up the Right Tree: Incorporating Biological Taxonomies into Deep Bioacoustic Classifiers
abstract
Class imbalance in the training data hinders the generalization ability of machine listening systems. In the context of bioacoustics, this issue may be circumvented by aggregating species labels into super-groups of higher taxonomic rank: genus, family, order, and so forth. However, different applications of machine listening to wildlife monitoring may require different levels of granularity. This paper introduces TaxoNet, a deep neural network for structured classification of signals from living organisms. TaxoNet is trained as a multitask and multilabel model, following a new architectural principle in end-to-end learning named "hierarchical composition": shallow layers extract a shared representation to predict a root taxon, while deeper layers specialize recursively to lower-rank taxa. In this way, TaxoNet is capable of handling taxonomic uncertainty, out-of-vocabulary labels, and open-set deployment settings. An experimental benchmark on two new bioacoustic datasets (ANAFCC and BirdVox-14SD) leads to state-of-the-art results in bird species classification. Furthermore, on a task of coarse-grained classification, TaxoNet also outperforms a flat single-task model trained on aggregate labels.
Jason Cramer, Vincent Lostanlen, Andrew Farnsworth, Justin Salamon, Juan Pablo Bello
ICASSP5
2020 Learning the Helix Topology of Musical Pitch
abstract
To explain the consonance of octaves, music psychologists represent pitch as a helix where azimuth and axial coordinate correspond to pitch class and pitch height respectively. This article addresses the problem of discovering this helical structure from unlabeled audio data. We measure Pearson correlations in the constant-Q transform (CQT) domain to build a K-nearest neighbor graph between frequency subbands. Then, we run the Isomap manifold learning algorithm to represent this graph in a three-dimensional space in which straight lines approximate graph geodesics. Experiments on isolated musical notes demonstrate that the resulting manifold resembles a helix which makes a full turn at every octave. A circular shape is also found in English speech, but not in urban noise. We discuss the impact of various design choices on the visualization: instrumentarium, loudness mapping function, and number of neighbors K.
Vincent Lostanlen, Sripathi Sridhar, Brian McFee, Andrew Farnsworth, Juan Pablo Bello
ICASSP5
2020 Few-Shot Sound Event Detection
abstract
Locating perceptually similar sound events within a continuous recording is a common task for various audio applications. However, current tools require users to manually listen to and label all the locations of the sound events of interest, which is tedious and time-consuming. In this work, we (1) adapt state-of-the-art metric-based few-shot learning methods to automate the detection of similar-sounding events, requiring only one or few examples of the target event, (2) develop a method to automatically construct a partial set of labeled examples (negative samples) to reduce user labeling effort, and (3) develop an inference-time data augmentation method to increase detection accuracy. To validate our approach, we perform extensive comparative analysis of few-shot learning methods for the task of keyword detection in speech. We show that our approach successfully adapts closed-set few-shot learning approaches to an open-set sound event detection problem.
Yu Wang 0105, Justin Salamon, Nicholas J. Bryan, Juan Pablo Bello
ICASSP4
2019 Crowdsourcing Multi-label Audio Annotation Tasks with Citizen Scientists
abstract
Annotating rich audio data is an essential aspect of training and evaluating machine listening systems. We approach this task in the context of temporally-complex urban soundscapes, which require multiple labels to identify overlapping sound sources. Typically this work is crowdsourced, and previous studies have shown that workers can quickly label audio with binary annotation for single classes. However, this approach can be difficult to scale when multiple passes with different focus classes are required to annotate data with multiple labels. In citizen science, where tasks are often image-based, annotation efforts typically label multiple classes simultaneously in a single pass. This paper describes our data collection on the Zooniverse citizen science platform, comparing the efficiencies of different audio annotation strategies. We compared multiple-pass binary annotation, single-pass multi-label annotation, and a hybrid approach: hierarchical multi-pass multi-label annotation. We discuss our findings, which support using multi-label annotation, with reference to volunteer citizen scientists' motivations.
Mark Cartwright, Graham Dove, Ana Elisa Méndez Méndez, Juan Pablo Bello, Oded Nov
CHI4
2019 Look, Listen, and Learn More: Design Choices for Deep Audio Embeddings
abstract
A considerable challenge in applying deep learning to audio classification is the scarcity of labeled data. An increasingly popular solution is to learn deep audio embeddings from large audio collections and use them to train shallow classifiers using small labeled datasets. Look, Listen, and Learn (L3-Net) is an embedding trained through self-supervised learning of audio-visual correspondence in videos as opposed to other embeddings requiring labeled data. This framework has the potential to produce powerful out-of-the-box embeddings for downstream audio classification tasks, but has a number of unexplained design choices that may impact the embeddings’ behavior. In this paper we investigate how L3-Net design choices impact the performance of downstream audio classifiers trained with these embeddings. We show that audio-informed choices of input representation are important, and that using sufficient data for training the embedding is key. Surprisingly, we find that matching the content for training the embedding to the downstream task is not beneficial. Finally, we show that our best variant of the L3-Net embedding outperforms both the VGGish and SoundNet embeddings, while having fewer parameters and being trained on less data. Our implementation of the L3-Net embedding model as well as pre-trained models are made freely available online.
Jason Cramer, Ho-Hsiang Wu, Justin Salamon, Juan Pablo Bello
ICASSP4
2019 A Music Structure Informed Downbeat Tracking System Using Skip-chain Conditional Random Fields and Deep Learning
abstract
In recent years the task of downbeat tracking has received increasing attention and the state of the art has been improved with the introduction of deep learning methods. Among proposed solutions, existing systems exploit short-term musical rules as part of their language modelling. In this work we show in an oracle scenario how including longer-term musical rules, in particular music structure, can enhance downbeat estimation. We introduce a skip-chain conditional random field language model for downbeat tracking designed to include section information in an unified and flexible framework. We combine this model with a state-of-the-art convolutional-recurrent network and we contrast the system’s performance to the commonly used Bar Pointer model. Our experiments on the popular Beatles dataset show that incorporating structure information in the language model leads to more consistent and more robust downbeat estimations.
Magdalena Fuentes, Brian McFee, Hélène C. Crayencour, Slim Essid, Juan Pablo Bello
ICASSP5
2019 Neural Music Synthesis for Flexible Timbre Control
abstract
The recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process - creating audio samples from a score and instrument information - is modeled using generative neural networks. This paper describes a neural music synthesis model with flexible timbre controls, which consists of a recurrent neural network conditioned on a learned instrument embedding followed by a WaveNet vocoder. The learned embedding space successfully captures the diverse variations in timbres within a large dataset and enables timbre control and morphing by interpolating between instruments in the embedding space. The synthesis quality is evaluated both numerically and perceptually, and an interactive web demo is presented.
Jong Wook Kim, Rachel M. Bittner, Aparna Kumar, Juan Pablo Bello
ICASSP4
2019 Active Learning for Efficient Audio Annotation and Classification with a Large Amount of Unlabeled Data
abstract
There are many sound classification problems that have target classes which are rare or unique to the context of the problem. For these problems, existing data sets are not sufficient and we must create new problem-specific datasets to train classification models. However, annotating a new dataset for every new problem is costly. Active learning could potentially reduce this annotation cost, but it has been understudied in the context of audio annotation. In this work, we investigate active learning to reduce the annotation cost of a sound classification dataset unique to a particular problem. We evaluate three certainty-based active learning query strategies and propose a new strategy: alternating confidence sampling. Using this strategy, we demonstrate reduced annotation costs when actively training models with both experts and non-experts, and we perform a qualitative analysis on 20k unlabeled recordings to show our approach results in a model that generalizes well to unseen data.
Yu Wang 0105, Ana Elisa Méndez Méndez, Mark Cartwright, Juan Pablo Bello
ICASSP4
2019 Per-Channel Energy Normalization: Why and How
abstract
In the context of automatic speech recognition and acoustic event detection, an adaptive procedure named per-channel energy normalization (PCEN) has recently shown to outperform the pointwise logarithm of mel-frequency spectrogram (logmelspec) as an acoustic frontend. This letter investigates the adequacy of PCEN for spectrogram-based pattern recognition in far-field noisy recordings, both from theoretical and practical standpoints. First, we apply PCEN on various datasets of natural acoustic environments and find empirically that it Gaussianizes distributions of magnitudes while decorrelating frequency bands. Second, we describe the asymptotic regimes of each component in PCEN: temporal integration, gain control, and dynamic range compression. Third, we give practical advice for adapting PCEN parameters to the temporal properties of the noise to be mitigated, the signal to be enhanced, and the choice of time-frequency representation. As it converts a large class of real-world soundscapes into additive white Gaussian noise, PCEN is a computationally efficient frontend for robust detection and classification of acoustic events in heterogeneous environments.
Vincent Lostanlen, Justin Salamon, Mark Cartwright, Brian McFee, Andrew Farnsworth, Steve Kelling, Juan Pablo Bello
IEEE Signal Process. Lett.7
2018 Investigating the Effect of Sound-Event Loudness on Crowdsourced Audio Annotations
abstract
Audio annotation is an important step in developing machine-listening systems. It is also a time consuming process, which has motivated investigators to crowdsource audio annotations. However, there are many factors that affect annotations, many of which have not been adequately investigated. In previous work, we investigated the effects of visualization aids and sound scene complexity on the quality of crowdsourced sound-event annotations. In this paper, we extend that work by investigating the effect of sound-event loudness on both sound-event source annotations and sound-event proximity annotations. We find that the sound class, loudness, and annotator bias affect how listeners annotate proximity. We also find that loudness affects recall more than precision and that the strengths of these effects are strongly influenced by the sound class. These findings are not only important for designing effective audio annotation processes, but also for effectively training and evaluating machine-listening systems.
Mark Cartwright, Justin Salamon, Ayanna Seals, Oded Nov, Juan Pablo Bello
ICASSP5
2018 Crepe: A Convolutional Representation for Pitch Estimation
abstract
The task of estimating the fundamental frequency of a monophonic sound recording, also known as pitch tracking, is fundamental to audio processing with multiple applications in speech processing and music information retrieval. To date, the best performing techniques, such as the pYIN algorithm, are based on a combination of DSP pipelines and heuristics. While such techniques perform very well on average, there remain many cases in which they fail to correctly estimate the pitch. In this paper, we propose a data-driven pitch tracking algorithm, CREPE, which is based on a deep convolutional neural network that operates directly on the time-domain waveform. We show that the proposed model produces state-of-the-art results, performing equally or better than pYIN. Furthermore, we evaluate the model's generalizability in terms of noise robustness. A pre-trained version of CREPE is made freely available as an open-source Python module for easy application.
Jong Wook Kim, Justin Salamon, Peter Li, Juan Pablo Bello
ICASSP4
2018 Birdvox-Full-Night: A Dataset and Benchmark for Avian Flight Call Detection
abstract
This article addresses the automatic detection of vocal, nocturnally migrating birds from a network of acoustic sensors. Thus far, owing to the lack of annotated continuous recordings, existing methods had been benchmarked in a binary classification setting (presence vs. absence). Instead, with the aim of comparing them in event detection, we release BirdVox-full-night, a dataset of 62 hours of audio comprising 35402 flight calls of nocturnally migrating birds, as recorded from 6 sensors. We find a large performance gap between energy-based detection functions and data-driven machine listening. The best model is a deep convolutional neural network trained with data augmentation. We correlate recall with the density of flight calls over time and frequency and identify the main causes of false alarm.
Vincent Lostanlen, Justin Salamon, Andrew Farnsworth, Steve Kelling, Juan Pablo Bello
ICASSP5
2018 Adaptive Pooling Operators for Weakly Labeled Sound Event Detection
abstract
Sound event detection (SED) methods are tasked with labeling segments of audio recordings by the presence of active sound sources. SED is typically posed as a supervised machine learning problem, requiring strong annotations for the presence or absence of each sound source at every time instant within the recording. However, strong annotations of this type are both labor- and cost-intensive for human annotators to produce, which limits the practical scalability of SED methods. In this paper, we treat SED as a multiple instance learning (MIL) problem, where training labels are static over a short excerpt, indicating the presence or absence of sound sources but not their temporal locality. The models, however, must still produce temporally dynamic predictions, which must be aggregated (pooled) when comparing against static labels during training. To facilitate this aggregation, we develop a family of adaptive pooling operators - referred to as autopool - which smoothly interpolate between common pooling operators, such as min-, max-, or average-pooling, and automatically adapt to the characteristics of the sound sources in question. We evaluate the proposed pooling operators on three datasets, and demonstrate that in each case, the proposed methods outperform nonadaptive pooling operators for static prediction, and nearly match the performance of models trained with strong, dynamic annotations. The proposed method is evaluated in conjunction with convolutional neural networks, but can be readily applied to any differentiable model for time-series label prediction. While this paper focuses on SED applications, the proposed methods are general, and could be applied widely to MIL problems in any domain.
Brian McFee, Justin Salamon, Juan Pablo Bello
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Pitch contour tracking in music using Harmonic Locked Loops
abstract
We present a novel time-domain pitch contour tracking algorithm based on Harmonic Locked Loops, which differs from existing method in terms of its approach, resolution, timbre information and speed. In addition to estimating pitch contours, the proposed method computes the amplitude of each harmonic over time, expanding the potential set of features that can be used for higher level tasks such as melody extraction. The method is tested against ground truth melody pitch annotations from publicly available datasets, and we show that contour recall is improved compared with a state of the art approach.
Rachel M. Bittner, Avery Wang, Juan Pablo Bello
ICASSP3
2017 Towards the characterization of singing styles in world music
abstract
In this paper we focus on the characterization of singing styles in world music.We develop a set of contour features capturing pitch structure and melodic embellishments.Using these features we train a binary classifier to distinguish vocal from non-vocal contours and learn a dictionary of singing style elements.Each contour is mapped to the dictionary elements and each recording is summarized as the histogram of its contour mappings.We use K-means clustering on the recording representations as a proxy for singing style similarity.We observe clusters distinguished by characteristic uses of singing techniques such as vibrato and melisma.Recordings that are clustered together are often from neighbouring countries or exhibit aspects of language and cultural proximity.Studying singing particularities in this comparative manner can contribute to understanding the interaction and exchange between world music styles.
Maria Panteli, Rachel M. Bittner, Juan Pablo Bello, Simon Dixon
ICASSP3
2017 Fusing shallow and deep learning for bioacoustic bird species classification
abstract
Automated classification of organisms to species based on their vocalizations would contribute tremendously to abilities to monitor biodiversity, with a wide range of applications in the field of ecology. In particular, automated classification of migrating birds' flight calls could yield new biological insights and conservation applications for birds that vocalize during migration. In this paper we explore state-of-the-art classification techniques for large-vocabulary bird species classification from flight calls. In particular, we contrast a “shallow learning” approach based on unsupervised dictionary learning with a deep convolutional neural network combined with data augmentation. We show that the two models perform comparably on a dataset of 5428 flight calls spanning 43 different species, with both significantly outperforming an MFCC baseline. Finally, we show that by combining the models using a simple late-fusion approach we can further improve the results, obtaining a state-of-the-art classification accuracy of 0.96.
Justin Salamon, Juan Pablo Bello, Andrew Farnsworth, Steve Kelling
ICASSP2
2017 Seeing Sound: Investigating the Effects of Visualizations and Complexity on Crowdsourced Audio Annotations
abstract
Audio annotation is key to developing machine-listening systems; yet, effective ways to accurately and rapidly obtain crowdsourced audio annotations is understudied. In this work, we seek to quantify the reliability/redundancy trade-off in crowdsourced soundscape annotation, investigate how visualizations affect accuracy and efficiency, and characterize how performance varies as a function of audio characteristics. Using a controlled experiment, we varied sound visualizations and the complexity of soundscapes presented to human annotators. Results show that more complex audio scenes result in lower annotator agreement, and spectrogram visualizations are superior in producing higher quality annotations at lower cost of time and human labor. We also found recall is more affected than precision by soundscape complexity, and mistakes can be often attributed to certain sound event characteristics. These findings have implications not only for how we should design annotation tasks and interfaces for audio data, but also how we train and evaluate machine-listening systems.
Mark Cartwright, Ayanna Seals, Justin Salamon, Alex C. Williams, Stefanie Mikloska, Duncan MacConnell, Edith Law, Juan Pablo Bello, Oded Nov
Proc. ACM Hum. Comput. Interact.8
2017 Deep Convolutional Neural Networks and Data Augmentation for Environmental Sound Classification
abstract
The ability of deep convolutional neural networks (CNNs) to learn discriminative spectro-temporal patterns makes them well suited to environmental sound classification. However, the relative scarcity of labeled data has impeded the exploitation of this family of high-capacity models. This study has two primary contributions: first, we propose a deep CNN architecture for environmental sound classification. Second, we propose the use of audio data augmentation for overcoming the problem of data scarcity and explore the influence of different augmentations on the performance of the proposed CNN architecture. Combined with data augmentation, the proposed model produces state-of-the-art results for environmental sound classification. We show that the improved performance stems from the combination of a deep, high-capacity model and an augmented training set: this combination outperforms both the proposed CNN without augmentation and a “shallow” dictionary learning model with augmentation. Finally, we examine the influence of each augmentation on the model's classification accuracy for each class, and observe that the accuracy for each class is influenced differently by each augmentation, suggesting that the performance of the model could be improved further by applying class-conditional data augmentation.
Justin Salamon, Juan Pablo Bello
IEEE Signal Process. Lett.2
2017 Robust Downbeat Tracking Using an Ensemble of Convolutional Networks
abstract
In this paper, we present a novel state-of-the-art system for automatic downbeat tracking from music signals. The audio signal is first segmented in frames which are synchronized at the tatum level of the music. We then extract different kind of features based on harmony, melody, rhythm, and bass content to feed convolutional neural networks that are adapted to take advantage of the characteristics of each feature. This ensemble of neural networks is combined to obtain one downbeat likelihood per tatum. The downbeat sequence is finally decoded with a flexible and efficient temporal model which takes advantage of the assumed metrical continuity of a song. We then perform an evaluation of our system on a large base of nine datasets, compare its performance to four other published algorithms and obtain a significant increase of 16.8% points compared to the second-best system, for altogether a moderate cost in test and training. The influence of each step of the method is studied to show its strengths and shortcomings.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Introduction to the Special Section on Sound Scene and Event Analysis
abstract
The papers in this special section are devoted to the growing field of acoustic scene classification and acoustic event recognition. Machine listening systems still have difficulties to reach the ability of human listeners in the analysis of realistic acoustic scenes. If sustained research efforts have been made for decades in speech recognition, speaker identification and to a lesser extent in music information retrieval, the analysis of other types of sounds, such as environmental sounds, is the subject of growing interest from the community and is targeting an ever increasing set of audio categories. This problem appears to be particularly challenging due to the large variety of potential sound sources in the scene, which may in addition have highly different acoustic characteristics, especially in bioacoustics. Furthermore, in realistic environments, multiple sources are often present simultaneously, and in reverberant conditions.
Gaël Richard, Tuomas Virtanen, Juan Pablo Bello, Nobutaka Ono, Hervé Glotin
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Feature adapted convolutional neural networks for downbeat tracking
abstract
We define a novel system for the automatic estimation of downbeat positions from audio music signals. New rhythm and melodic features are introduced and feature adapted convolutional neural networks are used to take advantage of their specificity. Indeed, invariance to melody transposition, chroma data augmentation and length-specific rhythmic patterns prove to be useful to learn downbeat likelihood. After the data is segmented in tatums, complementary features related to melody, rhythm and harmony are extracted and the likelihood of a tatum being at a downbeat position is computed with the aforementioned neural networks. The downbeat sequence is then extracted with a flexible temporal hidden Markov model. We then show the efficiency and robustness of our approach with a comparative evaluation conducted on 9 datasets.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
ICASSP2
2015 Downbeat tracking with multiple features and deep neural networks
abstract
In this paper, we introduce a novel method for the automatic estimation of downbeat positions from music signals. Our system relies on the computation of musically inspired features capturing important aspects of music such as timbre, harmony, rhythmic patterns, or local similarities in both timbre and harmony. It then uses several independent deep neural networks to learn higher-level representations. The downbeat sequences are finally obtained thanks to a temporal decoding step based on the Viterbi algorithm. The comparative evaluation conducted on varied datasets demonstrates the efficiency and robustness across different music styles of our approach.
Simon Durand, Juan Pablo Bello, Bertrand David 0002, Gaël Richard
ICASSP2
2015 Unsupervised feature learning for urban sound classification
abstract
Recent studies have demonstrated the potential of unsupervised feature learning for sound classification. In this paper we further explore the application of the spherical k-means algorithm for feature learning from audio signals, here in the domain of urban sound classification. Spherical k-means is a relatively simple technique that has recently been shown to be competitive with other more complex and time consuming approaches. We study how different parts of the processing pipeline influence performance, taking into account the specificities of the urban sonic environment. We evaluate our approach on the largest public dataset of urban sound sources available for research, and compare it to a baseline system based on MFCCs. We show that feature learning can outperform the baseline approach by configuring it to capture the temporal dynamics of urban sources. The results are complemented with error analysis and some proposals for future research.
Justin Salamon, Juan Pablo Bello
ICASSP2
2014 From music audio to chord tablature: Teaching deep convolutional networks toplay guitar
abstract
Automatic chord recognition is conventionally tackled as a general music audition task, where the desired output is a time-aligned sequence of discrete chord symbols, e.g. CMaj7, Esus2, etc. In practice, however, this presents two related challenges: one, the act of decoding a given chord sequence requires that the musician knows both the notes in the chord and how to play them on some instrument; and two, chord labeling systems do not degrade gracefully for users without significant musical training. Alternatively, we address both challenges by modeling the physical constraints of a guitar to produce human-readable representations of music audio, i.e guitar tablature via a deep convolutional network. Through training and evaluation as a standard chord recognition system, the model is able to yield representations that require minimal prior knowledge to interpret, while maintaining respectable performance compared to the state of the art.
Eric J. Humphrey, Juan Pablo Bello
ICASSP2
2014 Music segment similarity using 2D-Fourier Magnitude Coefficients
abstract
Music segmentation is the task of automatically identifying the different segments of a piece. In this work we present a novel approach to cluster the musical segments based on their acoustic similarity by using 2D-Fourier Magnitude Coefficients (2D-FMCs). These coefficients, computed from a chroma representation, significantly simplify the problem of clustering the different segments since they are key transposition and phase shift invariant. We explore various strategies to obtain the 2D-FMC patches that represent entire segments and apply k-means to label them. Finally, we discuss possible ways of estimating k and compare our competitive results with the current state of the art.
Oriol Nieto, Juan Pablo Bello
ICASSP2
2014 A Dataset and Taxonomy for Urban Sound Research
abstract
Automatic urban sound classification is a growing area of research with applications in multimedia retrieval and urban informatics. In this paper we identify two main barriers to research in this area - the lack of a common taxonomy and the scarceness of large, real-world, annotated data. To address these issues we present a taxonomy of urban sounds and a new dataset, UrbanSound, containing 27 hours of audio with 18.5 hours of annotated sound event occurrences across 10 sound classes. The challenges presented by the new dataset are studied through a series of experiments using a baseline classification system.
Justin Salamon, Christopher Jacoby, Juan Pablo Bello
ACM Multimedia3
2014 On the Relative Importance of Individual Components of Chord Recognition Systems
abstract
Most chord recognition systems share a common architecture comprising two main stages: feature extraction and pattern matching, and two optional sub stages: pre-filtering and post-filtering. Understanding the interaction between these basic components is very important not only for achieving optimal performance, but also for assessing the potential and limitations of the system. Unfortunately, there are no studies that sufficiently evaluate the effects of the different approaches to each processing step and the interactions between these steps. In this paper we attempt to remedy this deficiency by performing a systematic evaluation encompassing a wide variety of techniques used for each processing step. In our study we find that filtering has a significant impact on performance, but providing musical context information in the transition matrix is rendered moot by the need to enforce continuity in the estimations. We discovered that the benefits of using complex chord models can be largely offset by an appropriate choice of features. In addition, the initial performance gap between different features were not fully compensated by any subsequent processing stages.
Taemin Cho, Juan Pablo Bello
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Feature learning and deep architectures: new directions for music informatics
Eric J. Humphrey, Juan Pablo Bello, Yann LeCun
J. Intell. Inf. Syst.2
2012 Learning a robust Tonnetz-space transform for automatic chord recognition
abstract
Temporal pitch class profiles - commonly referred to as a chromagrams - are the de facto standard signal representation for content-based methods of musical harmonic analysis, despite exhibiting a set of practical difficulties. Here, we present a novel, data-driven approach to learning a robust function that projects audio data into Tonnetz-space, a geometric representation of equal-tempered pitch intervals grounded in music theory. We apply this representation to automatic chord recognition and show that our approach out-performs the classification accuracy of previous chroma representations, while providing a mid-level feature space that circumvents challenges inherent to chroma.
Eric J. Humphrey, Taemin Cho, Juan Pablo Bello
ICASSP3
2012 A Minimum Frame Error Criterion for Hidden Markov Model Training
abstract
Hidden Markov models (HMM) have been widely studied and applied over decades. The standard supervised learning method for HMM is maximum likelihood estimation (MLE) which maximizes the joint probability of training data. However, the most natural way of training would be finding the parameters that directly minimize the error rate of a given training set. In this article, we propose a novel learning method that minimizes the number of incorrectly decoded labels frame-wise. To do this, we construct a smooth function that is arbitrarily close to the exact frame error rate and minimize it directly using a gradient-based optimization algorithm. The proposed approach is intuitive and simple. We applied our method to the task of chord recognition in music, and the results show that it performs better than Maximum Likelihood Estimation and Minimum Classification Error.
Taemin Cho, Kibeom Kim, Juan Pablo Bello
ICMLA (2)3
2012 Rethinking Automatic Chord Recognition with Convolutional Neural Networks
abstract
Despite early success in automatic chord recognition, recent efforts are yielding diminishing returns while basically iterating over the same fundamental approach. Here, we abandon typical conventions and adopt a different perspective of the problem, where several seconds of pitch spectra are classified directly by a convolutional neural network. Using labeled data to train the system in a supervised manner, we achieve state of the art performance through this initial effort in an otherwise unexplored area. Subsequent error analysis provides insight into potential areas of improvement, and this approach to chord recognition shows promise for future harmonic analysis systems.
Eric J. Humphrey, Juan Pablo Bello
ICMLA (2)2
2011 Measuring Structural Similarity in Music
abstract
This paper presents a novel method for measuring the structural similarity between music recordings. It uses recurrence plot analysis to characterize patterns of repetition in the feature sequence, and the normalized compression distance, a practical approximation of the joint Kolmogorov complexity, to measure the pairwise similarity between the plots. By measuring the distance between intermediate representations of signal structure, the proposed method departs from common approaches to music structure analysis which assume a block-based model of music, and thus concentrate on segmenting and clustering sections. The approach ensures that global structure is consistently and robustly characterized in the presence of tempo, instrumentation, and key changes, while the used metric provides a simple to compute, versatile and robust alternative to common approaches in music similarity research. Finally, experimental results demonstrate success at characterizing similarity, while contributing an optimal parameterization of the proposed approach.
Juan Pablo Bello
IEEE Trans. Speech Audio Process.1
2007 Automatic Rhythm Modification of Drum Loops
abstract
We propose a novel system able to modify the rhythm of a given drum loop, known as the original, to match the rhythmic pattern of a second loop, known as the model. Our approach is fully automated, thus eliminating the need for MIDI sequencing. The presented methodology combines standard and state-of-the-art techniques for the segmentation and classification of drum sounds, the matching of drum sequences, and the transformation of the original loop. We discuss the advantages and disadvantages of the proposed approach and provide links to examples for the qualitative evaluation of the system's output
Emmanuel Ravelli, Juan Pablo Bello, Mark B. Sandler
IEEE Signal Process. Lett.2
2006 Drum Sound Analysis for the Manipulation of Rhythm in Drum Loops
abstract
This paper addresses the issue of drum sound classification in the context of automatic rhythm modification of drum loops. The proposed method segments the signal using an onset detection algorithm, characterises segmented sounds using a spectral feature set, and classifies them using k-means clustering. We propose a simple taxonomy for the grouping of different instrumental sounds under a few utilitarian labels. Results demonstrate the adequacy of our proposed taxonomy while showing that our classification approach outperforms commonly-used supervised learning techniques
Juan Pablo Bello, Emmanuel Ravelli, Mark B. Sandler
ICASSP (5)1
2006 Automatic Piano Transcription Using Frequency and Time-Domain Information
abstract
The aim of this paper is to propose solutions to some problems that arise in automatic polyphonic transcription of recorded piano music. First, we propose a method that groups spectral information in the frequency-domain and uses a rule-based framework to deal with the known problems of polyphony and harmonicity. Then, we present a novel method for multipitch-estimation that uses both frequency and time-domain information. It assumes signal segments to be the linearly weighted sum of waveforms in a database of individual piano notes. We propose a solution to the problem of generating those waveforms, by using the frequency-domain approach. We show that accurate time-domain transcription can be achieved given an adequate estimation of the database. This suggests an alternative to common frequency-domain approaches that does not require any prior training on a separate database of isolated notes
Juan Pablo Bello, Laurent Daudet, Mark B. Sandler
IEEE Trans. Speech Audio Process.1
2005 A Tutorial on Onset Detection in Music Signals
abstract
Note onset detection and localization is useful in a number of analysis and indexing techniques for musical signals. The usual way to detect onsets is to look for "transient" regions in the signal, a notion that leads to many definitions: a sudden burst of energy, a change in the short-time spectrum of the signal or in the statistical properties, etc. The goal of this paper is to review, categorize, and compare some of the most commonly used techniques for onset detection, and to present possible enhancements. We discuss methods based on the use of explicitly predefined signal features: the signal's amplitude envelope, spectral magnitudes and phases, time-frequency representations; and methods based on probabilistic signal models: model-based change point detection, surprise signals, etc. Using a choice of test cases, we provide some guidelines for choosing the appropriate method for a given application.
Juan Pablo Bello, Laurent Daudet, Samer A. Abdallah, Chris Duxbury, Mike E. Davies 0001, Mark B. Sandler
IEEE Trans. Speech Audio Process.1
2004 On the use of phase and energy for musical onset detection in the complex domain
abstract
We present a study on the combined use of energy and phase information for the detection of onsets in musical signals. The resulting method improves upon both energy-based and phase-based approaches. The detection function, generated from the analysis of the signal in the complex frequency domain is sharp at the position of onsets and smooth everywhere else. Results on a database of recordings show high detection rates for low rates of errors. The approach is more robust than its predecessors both theoretically and practically.
Juan Pablo Bello, Chris Duxbury, Mike E. Davies 0001, Mark B. Sandler
IEEE Signal Process. Lett.1
2003 Phase-based note onset detection for music signals
abstract
Note onsets mark the beginning of attack transients, short areas of a note containing rapid changes of the signal spectral content. Detecting onsets is not trivial, especially when analysing complex mixtures. Applications for note onset detection systems include time stretching, audio coding and synthesis. An alternative to standard energy-based onset detection is proposed by using phase information. It is suggested that by observing the frame-by-frame distribution of differential angles, the precise moment when onsets occur can be detected with accuracy. Statistical measures are used to build the detection function. The system is tested and tuned on a database of complex recordings.
Juan Pablo Bello, Mark B. Sandler
ICASSP (5)1
2002 Automatic Music Transcription and Audio Source Separation
abstract
In this article, we give an overview of a range of approaches to the analysis and separation of musical audio. In particular, we consider the problems of automatic music transcription and audio source separation, which are of particular interest to our group. Monophonic music transcription, where a single note is present at one time, can be tackled using an autocorrelation-based method. For polyphonic music transcription, with several notes at any time, other approaches can be used, such as a blackboard model or a multiple-cause/sparse coding method. The latter is based on ideas and methods related to independent component analysis (ICA), a method for sound source separation.
Mark D. Plumbley, Samer A. Abdallah, Juan Pablo Bello, Mike E. Davies 0001, Giuliano Monti, Mark B. Sandler
Cybern. Syst.3