Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Rachel M. Bittner

dblp:148/9664 · DBLP profile ↗
← Back
8ranked-venue papers
3as first author
3since 2021 · last 2024
0000-0001-7757-2232ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 2 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › music analysis
multimodal music analysis
0.812024
LLark: A Multimodal Instruction-Following Language Model for Music · ICML 2024
Audio and music processing › music information retrieval
music understanding
0.812024
LLark: A Multimodal Instruction-Following Language Model for Music · ICML 2024

Methods — techniques the papers use, named apart from their topics

pretrained generative music model · 1.5multimodal architecture · 1.5instruction tuning · 1.5
YearPublicationVenuePosition
2024 LLark: A Multimodal Instruction-Following Language Model for Music
abstract
Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal model for *music* understanding. We detail our process for dataset creation, which involves augmenting the annotations of diverse open-source music datasets and converting them to a unified instruction-tuning format. We propose a multimodal architecture for LLark, integrating a pretrained generative model for music with a pretrained language model. In evaluations on three types of tasks (music understanding, captioning, reasoning), we show that LLark matches or outperforms existing baselines in music understanding, and that humans show a high degree of agreement with its responses in captioning and reasoning tasks. LLark is trained entirely from open-source music data and models, and we make our training code available along with the release of this paper. Additional results and audio examples are at https://bit.ly/llark, and our source code is available at https://github.com/spotify-research/llark.
Josh Gardner 0001, Simon Durand, Daniel Stoller, Rachel M. Bittner
ICML4
2022 A Lightweight Instrument-Agnostic Model for Polyphonic Note Transcription and Multipitch Estimation
abstract
Automatic Music Transcription (AMT) has been recognized as a key enabling technology with a wide range of applications. Given the task’s complexity, best results have typically been reported for systems focusing on specific settings, e.g. instrument-specific systems tend to yield improved results over instrument-agnostic methods. Similarly, higher accuracy can be obtained when only estimating frame-wise f0values and neglecting the harder note event detection. Despite their high accuracy, such specialized systems often cannot be deployed in the real-world. Storage and network constraints prohibit the use of multiple specialized models, while memory and run-time constraints limit their complexity. In this paper, we propose a lightweight neural network for musical instrument transcription, which supports polyphonic outputs and generalizes to a wide variety of instruments (including vocals). Our model is trained to jointly predict frame-wise onsets, multipitch and note activations, and we experimentally show that this multi-output structure improves the resulting frame-level note accuracy. Despite its simplicity, benchmark results show our system’s note estimation to be substantially better than a comparable baseline, and its frame-level accuracy to be only marginally below those of specialized state-of-the-art AMT systems. With this work we hope to encourage the community to further investigate low-resource, instrument-agnostic AMT systems.
Rachel M. Bittner, Juan J. Bosch, David Rubinstein, Gabriel Meseguer-Brocal, Sebastian Ewert
ICASSP1
2022 Few-Shot Musical Source Separation
abstract
Deep learning-based approaches to musical source separation are often limited to the instrument classes that the models are trained on and do not generalize to separate unseen instruments. To address this, we propose a few-shot musical source separation paradigm. We condition a generic U-Net source separation model using few audio examples of the target instrument. We train a few-shot conditioning encoder jointly with the U-Net to encode the audio examples into a conditioning vector to configure the U-Net via feature-wise linear modulation (FiLM). We evaluate the trained models on real musical recordings in the MUSDB18 and MedleyDB datasets. We show that our proposed few-shot conditioning paradigm outperforms the base-line one-hot instrument-class conditioned model for both seen and unseen instruments. To extend the scope of our approach to a wider variety of real-world scenarios, we also experiment with different conditioning example characteristics, including examples from different recordings, with multiple sources, or negative conditioning examples.
Yu Wang 0105, Daniel Stoller, Rachel M. Bittner, Juan Pablo Bello
ICASSP3
2019 Neural Music Synthesis for Flexible Timbre Control
abstract
The recent success of raw audio waveform synthesis models like WaveNet motivates a new approach for music synthesis, in which the entire process - creating audio samples from a score and instrument information - is modeled using generative neural networks. This paper describes a neural music synthesis model with flexible timbre controls, which consists of a recurrent neural network conditioned on a learned instrument embedding followed by a WaveNet vocoder. The learned embedding space successfully captures the diverse variations in timbres within a large dataset and enables timbre control and morphing by interpolating between instruments in the embedding space. The synthesis quality is evaluated both numerically and perceptually, and an interactive web demo is presented.
Jong Wook Kim, Rachel M. Bittner, Aparna Kumar, Juan Pablo Bello
ICASSP2
2017 Pitch contour tracking in music using Harmonic Locked Loops
abstract
We present a novel time-domain pitch contour tracking algorithm based on Harmonic Locked Loops, which differs from existing method in terms of its approach, resolution, timbre information and speed. In addition to estimating pitch contours, the proposed method computes the amplitude of each harmonic over time, expanding the potential set of features that can be used for higher level tasks such as melody extraction. The method is tested against ground truth melody pitch annotations from publicly available datasets, and we show that contour recall is improved compared with a state of the art approach.
Rachel M. Bittner, Avery Wang, Juan Pablo Bello
ICASSP1
2017 Towards the characterization of singing styles in world music
abstract
In this paper we focus on the characterization of singing styles in world music.We develop a set of contour features capturing pitch structure and melodic embellishments.Using these features we train a binary classifier to distinguish vocal from non-vocal contours and learn a dictionary of singing style elements.Each contour is mapped to the dictionary elements and each recording is summarized as the histogram of its contour mappings.We use K-means clustering on the recording representations as a proxy for singing style similarity.We observe clusters distinguished by characteristic uses of singing techniques such as vibrato and melisma.Recordings that are clustered together are often from neighbouring countries or exhibit aspects of language and cultural proximity.Studying singing particularities in this comparative manner can contribute to understanding the interaction and exchange between world music styles.
Maria Panteli, Rachel M. Bittner, Juan Pablo Bello, Simon Dixon
ICASSP2
2015 Kernel Additive Modeling for interference reduction in multi-channel music recordings
abstract
When recording a live musical performance, the different voices, such as the instrument groups or soloists of an orchestra, are typically recorded in the same room simultaneously, with at least one microphone assigned to each voice. However, it is difficult to acoustically shield the microphones. In practice, each one contains interference from every other voice. In this paper, we aim to reduce these interferences in multi-channel recordings to recover only the isolated voices. Following the recently proposed Kernel Additive Modeling framework, we present a method that iteratively estimates both the power spectral density of each voice and the corresponding strength in each microphone signal. With this information, we build an optimal Wiener filter, strongly reducing interferences. The trade-off between distortion and separation can be controlled by the user through the number of iterations of the algorithm. Furthermore, we present a computationally effective approximation of the iterative procedure. Listening tests demonstrate the effectiveness of the method.
Thomas Prätzlich, Rachel M. Bittner, Antoine Liutkus, Meinard Müller
ICASSP2
2014 Signal processing methods for removing the effects of whole-body vibration upon speech
abstract
Humans may be exposed to whole-body vibration in environments where clear speech communications are crucial, particularly during the launch phase of space flight and in high-performance aircraft. Prior research has shown that high levels of vibration cause a decrease in speech intelligibility. However, the effects of whole-body vibration upon speech are not well understood, and no attempt has been made to restore speech distorted by whole-body vibration. In this paper, a model for speech during whole-body vibration is proposed and a method to remove its effect is described. The method presented reduces the perceptual effects of vibration, yields higher automatic speech recognition accuracy scores, and may significantly improve intelligibility. Possible applications include incorporation within spaceflight, aviation, or off-road vehicle radio-communication systems.
Rachel M. Bittner, Durand R. Begault
ICASSP1