EDBT 2026 Demo / reviewers in the wild / expert
Bryan Pardo
dblp:19/212
· DBLP profile ↗
64ranked-venue papers
2as first author
14since 2021 · last 2026
0000-0002-1427-6492ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 1 first-author · 10 since 2021Artificial intelligence and machine learning · 15 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 11 · 2 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Human-AI Interaction for Accessible CS Learning: Co-Designing AI with Blind and Visually Impaired LearnersabstractGenerative AI is increasingly used in learning environments, but most systems focus on generating responses rather than supporting interaction. This limitation is especially important for blind and visually impaired (BVI) learners, who rely on non-visual ways to navigate and understand programming tasks. We present findings from a co-design study with BVI high school students exploring expectations for AI assistants in expressive computer science learning. Participants described how AI could support accessible interaction, guide learning, and empower learner agency. We derive two design principles for inclusive AI assistants: (1) situated multimodal interaction and (2) scaffolded iterative interaction. These principles shift the focus from adapting outputs (e.g., audio) to designing interactions that support how learners navigate and engage in the CS learning context. Our work highlights the importance of interaction design in accessible AI and the value of involving BVI learners in the design of AI-powered educational technologies. Shi Ding, Jason Smith 0005, Kevin Gautier, Stephen Garrett, Brian Magerko, Jason Freeman 0001, Bryan Pardo, Stephanie Ludi, Taneisha Lee, Tom McKlin |
IDC | 7 |
| 2025 | Using Co-Design to Investigate Affordances of an Expressive CS Learning Environment for Students who are BVIabstractExpressive computer science (CS) learning environments teach coding through the creation of an artifact, such as audio or video output.EarSketch is an expressive CS learning environment designed to teach computing through music production, mixing and arranging sounds using code.In this paper, we explore the accessibility challenges of using EarSketch for learners who are Blind and Visually Impaired (BVI).We present key findings from co-design studies with teachers and students at an institution specializing in BVI education, focused on gathering both groups' unique perspectives about EarSketch's ability to support teachers' curricula, students' workflows using the system with accessibility software, and challenges faced by users who are BVI. Jason Smith 0005, Annie Chu, Noel Alben, Shi Ding, Kevin Gautier, Stephen Garrett, Brian Magerko, Jason Freeman 0001, Bryan Pardo, Stephanie Ludi, Taneisha Lee, Tom McKlin |
ASSETS | 9 |
| 2025 | Text2FX: Harnessing CLAP Embeddings for Text-Guided Audio EffectsabstractThis work introduces Text2FX, a method that leverages CLAP embeddings and differentiable digital signal processing to control audio effects, such as equalization and reverberation, using open-vocabulary natural language prompts (e.g., "make this sound in-your-face and bold"). Text2FX operates without retraining any models, relying instead on single-instance optimization within the existing embedding space, thus enabling a flexible, scalable approach to open-vocabulary sound transformations through interpretable and disentangled FX manipulation. We show that CLAP encodes valuable information for controlling audio effects and propose two optimization approaches using CLAP to map text to audio effect parameters. While we demonstrate with CLAP, this approach is applicable to any shared text-audio embedding space. Similarly, while we demonstrate with equalization and reverberation, any differentiable audio effect may be controlled. We conduct a listener study with diverse text prompts and source audio to evaluate the quality and alignment of these methods with human perception. Demos and code are available at anniejchu.github.io/text2fx Annie Chu, Patrick O'Reilly, Julia Barnett, Bryan Pardo |
ICASSP | 4 |
| 2025 | Sketch2Sound: Controllable Audio Generation via Time-Varying Signals and Sonic ImitationsabstractWe present Sketch2Sound, a generative audio model capable of creating high-quality sounds from a set of interpretable time-varying control signals: loudness, brightness, and pitch, as well as text prompts. Sketch2Sound can synthesize arbitrary sounds from sonic imitations (i.e., a vocal imitation or a reference sound-shape). Sketch2Sound can be implemented on top of any text-to-audio latent diffusion transformer (DiT), and requires only 40k steps of fine-tuning and a single linear layer per control, making it more lightweight than existing methods like ControlNet. To synthesize from sketchlike sonic imitations, we propose applying random median filters to the control signals during training, allowing Sketch2Sound to be prompted using controls with flexible levels of temporal specificity. We show that Sketch2Sound can synthesize sounds that follow the gist of input controls from a vocal imitation while retaining the adherence to an input text prompt and audio quality compared to a text-only baseline. Sketch2Sound allows sound artists to create sounds with the semantic flexibility of text prompts and the expressivity and precision of a sonic gesture or vocal imitation. Sound examples are available at https://hugofloresgarcia.art/sketch2sound/. Hugo Flores García, Oriol Nieto, Justin Salamon, Bryan Pardo, Prem Seetharaman |
ICASSP | 4 |
| 2025 | Code Drift: Towards Idempotent Neural Audio CodecsabstractNeural codecs have demonstrated strong performance in high-fidelity compression of audio signals at low bitrates. The token-based representations produced by these codecs have proven particularly useful for generative modeling. While much research has focused on improvements in compression ratio and perceptual transparency, recent works have largely overlooked another desirable codec property – idempotence, the stability of compressed outputs under multiple rounds of encoding. We find that state-of-the-art neural codecs exhibit varied degrees of idempotence, with some degrading audio outputs significantly after as few as three encodings. We investigate possible causes of low idempotence and devise a method for improving idempotence through fine-tuning a codec model. We then examine the effect of idempotence on a simple conditional generative modeling task, and find that increased idempotence can be achieved without negatively impacting downstream modeling performance – potentially extending the usefulness of neural codecs for practical file compression and iterative generative modeling workflows. Patrick O'Reilly, Prem Seetharaman, Jiaqu Su, Zeyu Jin, Bryan Pardo |
ICASSP | 5 |
| 2025 | WhAM: Towards A Translative Model of Sperm Whale VocalizationabstractSperm whales communicate in short sequences of clicks known as codas. We present WhAM (Whale Acoustics Model), the first transformer-based model capable of generating synthetic sperm whale codas from any audio prompt. WhAM is built by finetuning VampNet, a masked acoustic token model pretrained on musical audio, using 10k coda recordings collected over the past two decades. Through iterative masked token prediction, WhAM generates high-fidelity synthetic codas that preserve key acoustic features of the source recordings. We evaluate WhAM's synthetic codas using Fréchet Audio Distance and through perceptual studies with expert marine biologists. On downstream tasks including rhythm, social unit, and vowel classification, WhAM's learned representations achieve strong performance, despite being trained for generation rather than classification. Our code is available at https://github.com/Project-CETI/wham Orr Paradise, Liangyuan Chen, Pranav Muralikrishnan, Hugo Flores García, Bryan Pardo, Roee Diamant, David F. Gruber, Shane Gero, Shafi Goldwasser |
NeurIPS | 5 |
| 2024 | Crowdsourced and Automatic Speech Prominence EstimationabstractThe prominence of a spoken word is the degree to which an average native listener perceives the word as salient or emphasized relative to its context. Speech prominence estimation is the process of assigning a numeric value to the prominence of each word in an utterance. These prominence labels are useful for linguistic analysis, as well as training automated systems to perform emphasis-controlled text-to-speech or emotion recognition. Manually annotating prominence is time-consuming and expensive, which motivates the development of automated methods for speech prominence estimation. However, developing such an automated system using machine-learning methods requires human-annotated training data. Using our system for acquiring such human annotations, we collect and open-source crowd-sourced annotations of a portion of the LibriTTS dataset. We use these annotations as ground truth to train a neural speech prominence estimator that generalizes to unseen speakers, datasets, and speaking styles. We investigate design decisions for neural prominence estimation as well as how neural prominence estimation improves as a function of two key factors of annotation cost: dataset size and the number of annotations per utterance. Max Morrison, Pranav Pawar, Nathan Pruyne, Jennifer Cole 0001, Bryan Pardo |
ICASSP | 5 |
| 2024 | Maskmark: Robust Neuralwatermarking for Real and Synthetic SpeechabstractHigh-quality speech synthesis models may be used to spread misinformation or impersonate voices. Audio watermarking can combat misuse by embedding a traceable signature in generated audio. However, existing audio watermarks typically demonstrate robustness to only a small set of transformations of the watermarked audio. To address this, we propose MaskMark, a neural network-based digital audio watermarking technique optimized for speech. MaskMark embeds a secret key vector in audio via a multiplicative spectrogram mask, allowing the detection of watermarked speech segments even under substantial signal-processing or neural network-based transformations. Comparisons to a state-of-the-art baseline on natural and synthetic speech corpora and a human subjects evaluation demonstrate MaskMark’s superior robustness in detecting watermarked speech while maintaining high perceptual transparency. Patrick O'Reilly, Zeyu Jin, Jiaqi Su, Bryan Pardo |
ICASSP | 4 |
| 2024 | Fine-Grained and Interpretable Neural Speech Editing
Max Morrison, Cameron Churchwell, Nathan Pruyne, Bryan Pardo |
INTERSPEECH | 4 |
| 2022 | Improving Source Separation by Explicitly Modeling Dependencies between SourcesabstractWe propose a new method for training a supervised source separation system that aims to learn the interdependent relationships between all combinations of sources in a mixture. Rather than independently estimating each source from a mix, we reframe the source separation problem as an Orderless Neural Autoregressive Density Estimator (NADE), and estimate each source from both the mix and a random subset of the other sources. We adapt a standard source separation architecture, Demucs, with additional inputs for each individual source, in addition to the input mixture. We randomly mask these input sources during training so that the network learns the conditional dependencies between the sources. By pairing this training method with a blocked Gibbs sampling procedure at inference time, we demonstrate that the network can iteratively improve its separation performance by conditioning a source estimate on its earlier source estimates. Experiments on two source separation datasets show that training a Demucs model with an Orderless NADE approach and using Gibbs sampling (up to 512 steps) at inference time strongly outperforms a Demucs baseline that uses a standard regression loss and direct (one step) estimation of sources. Ethan Manilow, Curtis Hawthorne, Cheng-Zhi Anna Huang, Bryan Pardo, Jesse H. Engel |
ICASSP | 4 |
| 2022 | Source Separation By Steering Pretrained Music ModelsabstractWe showcase a method that repurposes deep models trained for music generation and music tagging for audio source separation, without any retraining. An audio generation model is conditioned on an input mixture, producing a latent encoding of the audio used to generate audio. This generated audio is fed to a pretrained music tagger that creates source labels. The cross-entropy loss between the tag distribution for the generated audio and a predefined distribution for an isolated source is used to guide gradient ascent in the (unchanging) latent space of the generative model. This system does not update the weights of the generative model or the tagger, and only relies on moving through the generative model’s latent space to produce separated sources. We use OpenAI’s Jukebox as the pretrained generative model, and we couple it with four kinds of pretrained music taggers (two architectures and two tagging datasets). Experimental results on two source separation datasets, show this approach can produce separation estimates for a wider variety of sources than any tested system. This work points to the vast and heretofore untapped potential of large pretrained music models for audio-to-audio tasks like source separation. Ethan Manilow, Patrick O'Reilly, Prem Seetharaman, Bryan Pardo |
ICASSP | 4 |
| 2022 | Effective and Inconspicuous Over-the-Air Adversarial Examples with Adaptive FilteringabstractWhile deep neural networks achieve state-of-the-art performance on many audio classification tasks, they are known to be vulnerable to adversarial examples - artificially-generated perturbations of natural instances that cause a network to make incorrect predictions. In this work we demonstrate a novel audio-domain adversarial attack that modifies benign audio using an interpretable and differentiable parametric transformation - adaptive filtering. Unlike existing state-of-the-art attacks, our proposed method does not require a complex optimization procedure or generative model, relying only on a simple variant of gradient descent to tune filter parameters. We demonstrate the effectiveness of our method by performing over-the-air attacks against a state-of-the-art speaker verification model and show that our attack is less conspicuous than an existing state-of-the-art attack while matching its effectiveness. Our results demonstrate the potential of transformations beyond direct waveform addition for concealing high-magnitude adversarial perturbations, allowing adversaries to attack more effectively in challenging, real-world settings. Patrick O'Reilly, Pranjal Awasthi, Aravindan Vijayaraghavan, Bryan Pardo |
ICASSP | 4 |
| 2022 | VoiceBlock: Privacy through Real-Time Adversarial Attacks with Audio-to-Audio ModelsabstractAs governments and corporations adopt deep learning systems to collect and analyze user-generated audio data, concerns about security and privacy naturally emerge in areas such as automatic speaker recognition. While audio adversarial examples offer one route to mislead or evade these invasive systems, they are typically crafted through time-intensive offline optimization, limiting their usefulness in streaming contexts. Inspired by architectures for audio-to-audio tasks such as denoising and speech enhancement, we propose a neural network model capable of adversarially modifying a user's audio stream in real-time. Our model learns to apply a time-varying finite impulse response (FIR) filter to outgoing audio, allowing for effective and inconspicuous perturbations on a small fixed delay suitable for streaming tasks. We demonstrate our model is highly effective at de-identifying user speech from speaker recognition and able to transfer to an unseen recognition system. We conduct a perceptual study and find that our method produces perturbations significantly less perceptible than baseline anonymization methods, when controlling for effectiveness. Finally, we provide an implementation of our model capable of running in real-time on a single CPU thread. Audio examples and code can be found at https://interactiveaudiolab.github.io/project/voiceblock.html. Patrick O'Reilly, Andreas Bugler, Keshav Bhandari, Max Morrison, Bryan Pardo |
NeurIPS | 5 |
| 2021 | Context-Aware Prosody Correction for Text-Based Speech EditingabstractText-based speech editors expedite the process of editing speech recordings by permitting editing via intuitive cut, copy, and paste operations on a speech transcript. A major drawback of current systems, however, is that edited recordings often sound unnatural because of prosody mismatches around edited regions. In our work, we propose a new context-aware method for more natural sounding text-based editing of speech. To do so, we 1) use a series of neural networks to generate salient prosody features that are dependent on the prosody of speech surrounding the edit and amenable to fine-grained user control 2) use the generated features to control a standard pitch-shift and time-stretch method and 3) apply a denoising neural network to remove artifacts induced by the signal manipulation to yield a high-fidelity result. We evaluate our approach using a subjective listening test, provide a detailed comparative analysis, and conclude several interesting insights. Max Morrison, Lucas Rencker, Zeyu Jin, Nicholas J. Bryan, Juan Pablo Cáceres, Bryan Pardo |
ICASSP | 6 |
| 2020 | Vroom!: A Search Engine for Sounds by Vocal Imitation QueriesabstractTraditional search through collections of audio recordings compares a text-based query to text metadata associated with each audio file and does not address the actual content of the audio. Text descriptions do not describe all aspects of the audio content in detail. Query by vocal imitation (QBV) is a kind of query by example that lets users imitate the content of the audio they seek, providing an alternative search method to traditional text search. Prior work proposed several neural networks, such as TL-IMINET, for QBV, however, previous systems have not been deployed in an actual search engine nor evaluated by real users. We have developed a state-of-the-art QBV system (Vroom!) and a baseline query-by-text search engine (TextSearch). We deployed both systems in an experimental framework to perform user experiments with Amazon Mechanical Turk (AMT) workers. Results showed that Vroom! received significantly higher search satisfaction ratings than TextSearch did for sound categories that were difficult for subjects to describe by text. Results also showed a better overall ease-of-use rating for Vroom! than TextSearch on the sound library used in our experiments. These findings suggest that QBV, as a complimentary search approach to existing text-based search, can improve both search results and user experience. Yichi Zhang 0008, Junbo Hu, Bryan Pardo, Zhiyao Duan |
CHIIR | 4 |
| 2020 | Simultaneous Separation and Transcription of Mixtures with Multiple Polyphonic and Percussive InstrumentsabstractWe present a single deep learning architecture that can both separate an audio recording of a musical mixture into constituent single-instrument recordings and transcribe these instruments into a human-readable format at the same time, learning a shared musical representation for both tasks. This novel architecture, which we call Cerberus, builds on the Chimera network for source separation by adding a third "head" for transcription. By training each head with different losses, we are able to jointly learn how to separate and transcribe up to five instruments with a single network. We show that separation and transcription are highly complementary with one another and when learned jointly, lead to Cerberus networks that are better at both separation and transcription and generalize better to unseen mixtures. Ethan Manilow, Prem Seetharaman, Bryan Pardo |
ICASSP | 3 |
| 2019 | VoiceAssist: Guiding Users to High-Quality Voice RecordingsabstractVoice recording is a challenging task with many pitfalls due to sub-par recording environments, mistakes in recording setup, microphone quality, etc. Newcomers to voice recording often have difficulty recording their voice, leading to recordings with low sound quality. Many amateur recordings of poor quality have two key problems: too much reverberation (echo), and too much background noise (e.g. fans, electronics, street noise). We present VoiceAssist, a system that helps inexperienced users produce high quality recordings by providing real-time visual feedback on audio quality. We integrate modern audio quality measures into an interactive human-machine feedback loop, so that the audio quality can be maximized at capture-time. We demonstrate the utility of this feedback for improving the recording quality with a user study. When presented with visual feedback about recording quality, users produced recordings that were strongly preferred by third-party listeners, when compared to recordings made without this feedback. Prem Seetharaman, Gautham J. Mysore, Bryan Pardo, Paris Smaragdis, Celso Gomes |
CHI | 3 |
| 2019 | Improving Content-based Audio Retrieval by Vocal Imitation FeedbackabstractContent-based audio retrieval including query-by-example (QBE) and query-by-vocal imitation (QBV) is useful when search-relevant text labels for the audio are unavailable, or text labels do not sufficiently narrow the search. However, a single query example may not provide sufficient information to ensure the target sound(s) in the database are the most highly ranked. In this paper, we adapt an existing model for generating audio embeddings to create a state-of-the-art similarity measure for audio QBE and QBV. We then propose a new method to update search results when top-ranked items are not relevant: The user provides an additional vocal imitation to illustrate what they do or do not want in the search results. This imitation may either be of some portion of the initial query example, or of a top-ranked (but incorrect) search result. Results show that adding vocal imitation feedback improves initial retrieval results by a statistically significant amount. Bongjun Kim, Bryan Pardo |
ICASSP | 2 |
| 2019 | Bootstrapping Single-channel Source Separation via Unsupervised Spatial Clustering on Stereo MixturesabstractSeparating an audio scene into isolated sources is a fundamental problem in computer audition, analogous to image segmentation in visual scene analysis. Source separation systems based on deep learning are currently the most successful approaches for solving the underdetermined separation problem, where there are more sources than channels. Such systems are normally trained on sound mixtures where the ground truth decomposition is already known. In this work, we use an unsupervised spatial source separation on stereo mixtures which generates initial decompositions of mixtures to train a deep learning source separation model. These estimated decompositions vary greatly in quality across the training mixtures. To overcome this, we weight the data during training using a confidence measure that assesses which mixtures or parts of mixtures are well-separated by the unsupervised algorithm. Once trained, the model can be applied to separate single-channel mixtures, where no source direction information is available. The idea is to use simple, low-level processing to separate sources in an unsupervised fashion, identify easy conditions, and then use that knowledge to bootstrap a (self-)supervised source separation model for difficult conditions. We also explore using the two approaches in an ensemble. Prem Seetharaman, Gordon Wichern, Jonathan Le Roux, Bryan Pardo |
ICASSP | 4 |
| 2019 | Multi-Resolution Common Fate TransformabstractThe multi-resolution common fate transform (MCFT) is an audio signal representation useful for representing mixtures of multiple audio signals that overlap in both time and frequency. The MCFT combines the invertibility of a state-of-the-art representation, the common fate transform (CFT), and the multi-resolution property of the cortical stage output of an auditory model. Since the MCFT is computed based on a fully invertible complex time-frequency representation, separation of audio sources with high time-frequency overlap may be performed directly in the MCFT domain, where there is less overlap between sources than in the time-frequency domain. The MCFT circumvents the resolution issue of the CFT by using a multi-resolution two-dimensional (2D) filter bank instead of fixed-size 2D windows. This enables higher quality separation without the need to hand-tune the window size to the specific case. In this work, we describe the MCFT, discuss the properties of the MCFT with the aid of illustrative examples, and provide definitions and objective measures for two desirable representation properties: separability of source signals andclusterabilityof components of each signal. The utility of the MCFT for source separation is illustrated by performing ideal masking on a comprehensive dataset of audio mixtures of musical tones played in unison, including audio samples from a wide pitch range and a variety of instruments/playing techniques. Results show that the ideal masks made in the MCFT domain yield better separability than those made in commonly used time-frequency signal representations as well as the CFT. The use of the MCFT also results in more reliable clusterability than the CFT in most cases. Fatemeh Pishdadian, Bryan Pardo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Siamese Style Convolutional Neural Networks for Sound Search by Vocal ImitationabstractConventional methods for finding audio in databases typically search text labels, rather than the audio itself. This can be problematic as labels may be missing, irrelevant to the audio content, or not known by users. Query by vocal imitation lets users query using vocal imitations instead. To do so, appropriate audio feature representations and effective similarity measures of imitations and original sounds must be developed. In this paper, we build upon our preliminary work to propose Siamese style convolutional neural networks to learn feature representations and similarity measures in a unified end-to-end training framework. Our Siamese architecture uses two convolutional neural networks to extract features, one from vocal imitations and the other from original sounds. The encoded features are then concatenated and fed into a fully connected network to estimate their similarity. We propose two versions of the system: IMINET is symmetric where the two encoders have an identical structure and are trained from scratch, while TL-IMINET is asymmetric and adopts the transfer learning idea by pretraining the two encoders from other relevant tasks: spoken language recognition for the imitation encoder and environmental sound classification for the original sound encoder. Experimental results show that both versions of the proposed system outperform a state-of-the-art system for sound search by vocal imitation, and the performance can be further improved when they are fused with the state of the art system. Results also show that transfer learning significantly improves the retrieval performance. This paper also provides insights to the proposed networks by visualizing and sonifying input patterns that maximize the activation of certain neurons in different layers. Yichi Zhang 0008, Bryan Pardo, Zhiyao Duan |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Crowdsourced Pairwise-Comparison for Source Separation EvaluationabstractAutomated objective methods of audio source separation evaluation are fast, cheap, and require little effort by the investigator. However, their output often correlates poorly with human quality assessments and typically require ground-truth (perfectly separated) signals to evaluate algorithm performance. Subjective multi-stimulus human ratings (e.g. MUSHRA) of audio quality are the gold standard for many tasks, but they are slow and require a great deal of effort to recruit participants and run listening tests. Recent work has shown that a crowdsourced multi-stimulus listening test can have results comparable to lab-based multi-stimulus tests. While these results are encouraging, MUSHRA multi-stimulus tests are limited to evaluating 12 or fewer stimuli, and they require ground-truth stimuli for reference. In this work, we evaluate a web-based pairwise-comparison listening approach that promises to speed and facilitate conducting listening tests, while also addressing some of the shortcomings of multi-stimulus tests. Using audio source separation quality as our evaluation task, we compare our web-based pairwise-comparison listening test to both web-based and lab-based multi-stimulus tests. We find that pairwise-comparison listening tests perform comparably to multi-stimulus tests, but without many of their shortcomings. Mark Cartwright, Bryan Pardo, Gautham J. Mysore |
ICASSP | 2 |
| 2018 | Blind Estimation of the Speech Transmission Index for Speech Quality PredictionabstractThe speech transmission index (STI) of a listening position within a given room indicates the quality and intelligibility of speech uttered in that room. The measure is very reliable for predicting speech intelligibility in many room conditions but requires an STI measurement of the impulse response for the room. We present a method for blindly estimating the STI without measuring or modeling the impulse response of the room using deep convolutional neural networks. Our model is trained entirely using simulated room impulse responses combined with clean speech examples from the DAPS dataset [1] and works directly on PCM audio. Our experiments show that our method predicts true STI with a high degree of accuracy - an average error of under 4%. It can also distinguish between different STI conditions to a level of granularity that is comparable to humans. Prem Seetharaman, Gautham J. Mysore, Paris Smaragdis, Bryan Pardo |
ICASSP | 4 |
| 2018 | An Overview of Lead and Accompaniment Separation in MusicabstractPopular music is often composed of an accompaniment and a lead component, the latter typically consisting of vocals. Filtering such mixtures to extract one or both components has many applications, such as automatic karaoke and remixing. This particular case of source separation yields very specific challenges and opportunities, including the particular complexity of musical structures, but also relevant prior knowledge coming from acoustics, musicology or sound engineering. Due to both its importance in applications and its challenging difficulty, lead and accompaniment separation has been a popular topic in signal processing for decades. In this article, we provide a comprehensive review of this research topic, organizing the different approaches according to whether they are model-based or data-centered. For model-based methods, we organize them according to whether they concentrate on the lead signal, the accompaniment, or both. For data-centered approaches, we discuss the particular difficulty of obtaining data for learning lead separation systems, and then review recent approaches, notably those based on deep learning. Finally, we discuss the delicate problem of evaluating the quality of music separation through adequate metrics and present the results of the largest evaluation, to-date, of lead and accompaniment separation systems. In conjunction with the above, a comprehensive list of references is provided, along with relevant pointers to available implementations and repositories. Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos I. Mimilakis, Derry Fitzgerald, Bryan Pardo |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | A Human-in-the-Loop System for Sound Event Detection and AnnotationabstractLabeling of audio events is essential for many tasks. However, finding sound events and labeling them within a long audio file is tedious and time-consuming. In cases where there is very little labeled data (e.g., a single labeled example), it is often not feasible to train an automatic labeler because many techniques (e.g., deep learning) require a large number of human-labeled training examples. Also, fully automated labeling may not show sufficient agreement with human labeling for many uses. To solve this issue, we present a human-in-the-loop sound labeling system that helps a user quickly label target sound events in a long audio. It lets a user reduce the time required to label a long audio file (e.g., 20 hours) containing target sounds that are sparsely distributed throughout the recording (10% or less of the audio contains the target) when there are too few labeled examples (e.g., one) to train a state-of-the-art machine audio labeling system. To evaluate the effectiveness of our tool, we performed a human-subject study. The results show that it helped participants label target sound events twice as fast as labeling them manually. In addition to measuring the overall performance of the proposed system, we also measure interaction overhead and machine accuracy, which are two key factors that determine the overall performance. The analysis shows that an ideal interface that does not have interaction overhead at all could speed labeling by as much as a factor of four. Bongjun Kim, Bryan Pardo |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2017 | A Multi-resolution approach to Common Fate-based audio separationabstractWe propose the Multi-resolution Common Fate Transform (MCFT), a signal representation that increases the separability of audio sources with significant energy overlap in the time-frequency domain. The MCFT combines the desirable features of two existing representations: the invertibility of the recently proposed Common Fate Transform (CFT) and the multi-resolution property of the cortical stage output of an auditory model. We compare the utility of the MCFT to the CFT by measuring the quality of source separation performed via ideal binary masking using each representation. Experiments on harmonic sounds with overlapping fundamental frequencies and different spectro-temporal modulation patterns show that ideal masks based on the MCFT yield better separation than those based on the CFT. Fatemeh Pishdadian, Bryan Pardo, Antoine Liutkus |
ICASSP | 2 |
| 2017 | I-SED: An Interactive Sound Event DetectorabstractTagging of sound events is essential in many research areas. However, finding sound events and labeling them within a long audio file is tedious and time-consuming. Building an automatic recognition system using machine learning techniques is often not feasible because it requires a large number of human-labeled training examples and fine tuning the model for a specific application. Fully automated labeling is also not reliable enough for all uses. We present I-SED, an interactive sound detection interface using a human-in-the-loop approach that lets a user reduce the time required to label audio that is tediously long (e.g. 20 hours) to do manually and has too few prior labeled examples (e.g. one) to train a state-of-the-art machine audio labeling system. We performed a human-subject study to validate its effectiveness and the results showed that our tool helped participants label all target sound events within a recording twice as fast as labeling them manually. Bongjun Kim, Bryan Pardo |
IUI | 2 |
| 2016 | An Approach to Audio-Only Editing for Visually Impaired SeniorsabstractOlder adults and people with vision impairments are increasingly using phones to receive audio-based information and want to publish content online but must use complex audio recording/editing tools that often rely on inaccessible graphical interfaces. This poster describes the design of an accessible audio-based interface for post-processing audio content created by visually impaired seniors. We conducted a diary study with five older adults with vision impairments to understand how to design a system that would allow them to edit content they record using an audio-only interface. Our findings can help inform the development of accessible audio-editing interfaces for people with vision impairments more broadly. Robin Brewer, Mark Cartwright, Aaron Karp, Bryan Pardo, Anne Marie Piper |
ASSETS | 4 |
| 2016 | Fast and easy crowdsourced perceptual audio evaluationabstractAutomated objective methods of audio evaluation are fast, cheap, and require little effort by the investigator. However, objective evaluation methods do not exist for the output of all audio processing algorithms, often have output that correlates poorly with human quality assessments, and require ground truth data in their calculation. Subjective human ratings of audio quality are the gold standard for many tasks, but are expensive, slow, and require a great deal of effort to recruit subjects and run listening tests. Moving listening tests from the lab to the micro-task labor market of Amazon Mechanical Turk speeds data collection and reduces investigator effort. However, it also reduces the amount of control investigators have over the testing environment, adding new variability and potential biases to the data. In this work, we compare multiple stimulus listening tests performed in a lab environment to multiple stimulus listening tests performed in web environment on a population drawn from Mechanical Turk. Mark Cartwright, Bryan Pardo, Gautham J. Mysore, Matthew Hoffman 0001 |
ICASSP | 2 |
| 2016 | SocialFX: Studying a Crowdsourced Folksonomy of Audio Effects TermsabstractWe present the analysis of crowdsourced studies into how a population of Amazon Mechanical Turk Workers describe three commonly used audio effects: equalization, reverberation, and dynamic range compression. We find three categories of words used to describe audio: ones that are generally used across effects, ones that tend towards a single effect, and ones that are exclusive to a single effect. We present select examples from these categories. We visualize and present an analysis of the shared descriptor space between audio effects. Data on the strength of association between words and effects is made available online for a set of 4297 words drawn from 1233 unique users for three effects (equalization, reverberation, compression). This dataset is an important step towards implementing of an end-to-end language-based audio production system, in which a user describes a creative goal, as they would to a professional audio engineer, and the system picks which audio effect to apply, as well as the setting of the audio effect. Taylor Zheng, Prem Seetharaman, Bryan Pardo |
ACM Multimedia | 3 |
| 2015 | VocalSketch: Vocally Imitating Audio ConceptsabstractA natural way of communicating an audio concept is to imitate it with one's voice. This creates an approximation of the imagined sound (e.g. a particular owl's hoot), much like how a visual sketch approximates a visual concept (e.g a drawing of the owl). If a machine could understand vocal imitations, users could communicate with software in this natural way, enabling new interactions (e.g. programming a music synthesizer by imitating the desired sound with one's voice). In this work, we collect thousands of crowd-sourced vocal imitations of a large set of diverse sounds, along with data on the crowd's ability to correctly label these vocal imitations. The resulting data set will help the research community understand which audio concepts can be effectively communicated with this approach. We have released the data set so the community can study the related issues and build systems that leverage vocal imitation as an interaction modality. Mark Cartwright, Bryan Pardo |
CHI | 2 |
| 2015 | A simple user interface system for recovering patterns repeating in time and frequency in mixtures of soundsabstractRepetition is a fundamental element in generating and perceiving structure in audio. Especially in music, structures tend to be composed of patterns that repeat through time (e.g., rhythmic elements in a musical accompaniment), and also frequency (e.g., different notes of the same instrument). The auditory system has the remarkable ability to parse such patterns by identifying repetitions within the audio mixture. On this basis, we propose a simple user interface system for recovering patterns repeating in time and frequency in mixtures of sounds. A user selects a region in the log-frequency spectrogram of an audio recording from which she/he wishes to recover a repeating pattern masked by an undesired element (e.g., a note masked by a cough). The selected region is then cross-correlated with the spectrogram to identify similar regions where the underlying pattern repeats. The identified regions are finally averaged over their repetitions and the repeating pattern is recovered. Zafar Rafii, Antoine Liutkus, Bryan Pardo |
ICASSP | 3 |
| 2014 | Adapting Collaborative Filtering to Personalized Audio ProductionabstractRecommending media objects to users typically requires users to rate existing media objects so as to understand their preferences. The number of ratings required to produce good suggestions can be reduced through collaborative filtering. Collaborative filtering is more difficult when prior users have not rated the same set of media objects as the current user or each other. In this work, we describe an approach to applying prior user data in a way that does not require users to rate the same media objects and that does not require imputation (estimation) of prior user ratings of objects they have not rated. This approach is applied to the problem of finding good equalizer settings for music audio and is shown to greatly reduce the number of ratings the current user must make to find a good equalization setting. Bongjun Kim, Bryan Pardo |
HCOMP | 2 |
| 2014 | A novel cepstral representation for timbre modeling of sound sources in polyphonic mixturesabstractWe propose a novel cepstral representation called the uniform discrete cepstrum (UDC) to represent the timbre of sound sources in a sound mixture. Different from ordinary cepstrum and MFCC which have to be calculated from the full magnitude spectrum of a source after source separation, UDC can be calculated directly from isolated spectral points that are likely to belong to the source in the mixture spectrum (e.g., non-overlapping harmonics of a harmonic source). Existing cepstral representations that have this property are discrete cepstrum and regularized discrete cepstrum, however, compared to the proposed UDC, they are not as effective and are more complex to compute. The key advantage of UDC is that it uses a more natural and locally adaptive regularizer to prevent it from overfitting the isolated spectral points. We derive the mathematical relations between these cepstral representations, and compare their timbre modeling performances in the task of instrument recognition in polyphonic audio mixtures. We show that UDC and its mel-scale variant MUDC significantly outperform all the other representations. Zhiyao Duan, Bryan Pardo, Laurent Daudet |
ICASSP | 2 |
| 2014 | Speeding Learning of Personalized Audio EqualizationabstractAudio equalizers (EQs) are perhaps the most commonly used tools used in audio production. The SocialEQ project is a web-based personalized audio equalization system that uses an alternative interface paradigm to the standard approach. Here, the user names a desired effect (e.g. Make the sound "warm") and teaches the tool (e.g. An equalizer) what settings make the sound embody the term. Social EQ typically requires 25 ratings to properly personalize the equalization settings. In this paper, we present three methods to improve the speed of generating personalized items (audio settings) so users can be provided personalized EQ curves after rating a much smaller number of examples. These methods can be adapted to any situation where collaborative filtering is desirable, the end products created for users are unique and comparable to each other, but prior users did not rate the same set of examples as the current user. Methods are tested on a data set of 1635 user sessions. Bongjun Kim, Bryan Pardo |
ICMLA | 2 |
| 2014 | MIXPLORATION: rethinking the audio mixer interfaceabstractA typical audio mixer interface consists of faders and knobs that control the amplitude level as well as processing (e.g. equalization, compression and reverberation) parameters of individual tracks. This interface, while widely used and effective for optimizing a mix, may not be the best interface to facilitate exploration of different mixing options. In this work, we rethink the mixer interface, describing an alternative interface for exploring the space of possible mixes of four audio tracks. In a user study with 24 participants, we compared the effectiveness of this interface to the traditional paradigm for exploring alternative mixes. In the study, users responded that the proposed alternative interface facilitated exploration and that they considered the process of rating mixes to be beneficial. Mark Cartwright, Bryan Pardo, Joshua D. Reiss |
IUI | 2 |
| 2014 | SynthAssist: an audio synthesizer programmed with vocal imitationabstractWhile programming an audio synthesizer can be difficult, if a user has a general idea of the sound they are trying to program, they may be able to imitate it with their voice. In this technical demonstration, we demonstrate SynthAssist, a system that allows the user to program an audio synthesizer using vocal imitation and interactive feedback. This system treats synthesizer programming as an audio information retrieval task. To account for the limitations of the human voice, it compares vocal imitations to synthesizer sounds by using both absolute and relative temporal shapes of relevant audio features, and it refines the query and feature weights using relevance feedback. Mark Cartwright, Bryan Pardo |
ACM Multimedia | 2 |
| 2014 | Crowdsourcing a Reverberation Descriptor MapabstractAudio production is central to every kind of media that involves sound, such as film, television, and music and involves transforming audio into a state ready for consumption by the public. One of the most commonly-used audio production tools is the reverberator. Current interfaces are often complex and hard-to-understand. We seek to simplify these interfaces by letting users communicate their audio production objective with descriptive language (e.g. "Make the drums sound bigger."). To achieve this goal, a system must be able to tell whether the stated goal is appropriate for the selected tool (e.g. making the violin warmer using a panning tool does not make sense). If the goal is appropriate for the tool, it must know what actions lead to the goal. Further, the tool should not impose a vocabulary on users, but rather understand the vocabulary users prefer. In this work, we describe SocialReverb, a project to crowdsource a vocabulary of audio descriptors that can be mapped onto concrete actions using a parametric reverberator. We deployed SocialReverb, on Mechanical Turk, where 513 unique users described 256 instances of reverberation using 2861 unique words. We used this data to build a concept map showing which words are popular descriptors, which ones map consistently to specific reverberation types, and which ones are synonyms. This promises to enable future interfaces that let the user communicate their production needs using natural language. Prem Seetharaman, Bryan Pardo |
ACM Multimedia | 2 |
| 2014 | Reverbalize: A Crowdsourced Reverberation ControllerabstractOne of the most commonly-used audio production tools is the reverberator. Reverberators apply subtle or large echo effects to sound and are typically used in commercial audio recordings. Current reverberator interfaces are often complex and hard-to-understand. In this work, we describe Reverbalize, a novel and easy-to-use interface for a reverberator. Reverbalize uses crowdsourced data to create a 2-dimensional map of adjectives used to describe reverberation (e.g. "underwater'). Adjacent words describe similar reverberation effects. Word size correlates with agreement for the definition of a word. To use Reverbalize, the user simply clicks on the descriptive adjective that best describes the desired effect. The tool modifies the sound accordingly. A text search box also lets the user type in the desired word. Prem Seetharaman, Bryan Pardo |
ACM Multimedia | 2 |
| 2014 | Multi-pitch Streaming of Harmonic Sound MixturesabstractMulti-pitch analysis of concurrent sound sources is an important but challenging problem. It requires estimating pitch values of all harmonic sources in individual frames and streaming the pitch estimates into trajectories, each of which corresponds to a source. We address the streaming problem for monophonic sound sources. We take the original audio, plus frame-level pitch estimates from any multi-pitch estimation algorithm as inputs, and output a pitch trajectory for each source. Our approach does not require pre-training of source models from isolated recordings. Instead, it casts the problem as a constrained clustering problem, where each cluster corresponds to a source. The clustering objective is to minimize the timbre inconsistency within each cluster. We explore different timbre features for music and speech. For music, harmonic structure and a newly proposed feature called uniform discrete cepstrum (UDC) are found effective; while for speech, MFCC and UDC works well. We also show that timbre-consistency is insufficient for effective streaming. Constraints are imposed on pairs of pitch estimates according to their time-frequency relationships. We propose a new constrained clustering algorithm that satisfies as many constraints as possible while optimizing the clustering objective. We compare the proposed approach with other state-of-the-art supervised and unsupervised multi-pitch streaming approaches that are specifically designed for music or speech. Better or comparable results are shown. Zhiyao Duan, Jinyu Han, Bryan Pardo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Combining rhythm-based and pitch-based methods for background and melody separationabstractMusical works are often composed of two characteristic components: the background (typically the musical accompaniment), which generally exhibits a strong rhythmic structure with distinctive repeating time elements, and the melody (typically the singing voice or a solo instrument), which generally exhibits a strong harmonic structure with a distinctive predominant pitch contour. Drawing from findings in cognitive psychology, we propose to investigate the simple combination of two dedicated approaches for separating those two components: a rhythm-based method that focuses on extracting the background via a rhythmic mask derived from identifying the repeating time elements in the mixture and a pitch-based method that focuses on extracting the melody via a harmonic mask derived from identifying the predominant pitch contour in the mixture. Evaluation on a data set of song clips showed that combining such two contrasting yet complementary methods can help to improve separation performance-from the point of view of both components-compared with using only one of those methods, and also compared with two other state-of-the-art approaches. Zafar Rafii, Zhiyao Duan, Bryan Pardo |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2013 | Online REPET-SIM for real-time speech enhancementabstractREPET-SIM is a generalization of the REpeating Pattern Extraction Technique (REPET) that uses a similarity matrix to separate the repeating background from the non-repeating foreground in a mixture. The method assumes that the background (typically the music accompaniment) is dense and low-ranked, while the foreground (typically the singing voice) is sparse and varied. While this assumption is often true for background music and foreground voice in musical mixtures, it also often holds for background noise and foreground speech in noisy mixtures. We therefore propose here to extend REPET-SIM for noise/speech segregation. In particular, given the low computational complexity of the algorithm, we show that the method can be easily implemented online for real-time processing. Evaluation on a data set of 10 stereo two-channel mixtures of speech and real-world background noise showed that this online REPET-SIM can be successfully applied for real-time speech enhancement, performing as well as different competitive methods. Zafar Rafii, Bryan Pardo |
ICASSP | 2 |
| 2013 | REpeating Pattern Extraction Technique (REPET): A Simple Method for Music/Voice SeparationabstractRepetition is a core principle in music. Many musical pieces are characterized by an underlying repeating structure over which varying elements are superimposed. This is especially true for pop songs where a singer often overlays varying vocals on a repeating accompaniment. On this basis, we present the REpeating Pattern Extraction Technique (REPET), a novel and simple approach for separating the repeating “background” from the non-repeating “foreground” in a mixture. The basic idea is to identify the periodically repeating segments in the audio, compare them to a repeating segment model derived from them, and extract the repeating patterns via time-frequency masking. Experiments on data sets of 1,000 song clips and 14 full-track real-world songs showed that this method can be successfully applied for music/voice separation, competing with two recent state-of-the-art approaches. Further experiments showed that REPET can also be used as a preprocessor to pitch detection algorithms to improve melody extraction. Zafar Rafii, Bryan Pardo |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Adaptive filtering for music/voice separation exploiting the repeating musical structureabstractThe separation of the lead vocals from the background accompaniment in audio recordings is a challenging task. Recently, an efficient method called REPET (REpeating Pattern Extraction Technique) has been proposed to extract the repeating background from the non-repeating foreground. While effective on individual sections of a song, REPET does not allow for variations in the background (e.g. verse vs. chorus), and is thus limited to short excerpts only. We overcome this limitation and generalize REPET to permit the processing of complete musical tracks. The proposed algorithm tracks the period of the repeating structure and computes local estimates of the background pattern. Separation is performed by soft time-frequency masking, based on the deviation between the current observation and the estimated background pattern. Evaluation on a dataset of 14 complete tracks shows that this method can perform at least as well as a recent competitive music/voice separation method, while being computationally efficient. Antoine Liutkus, Zafar Rafii, Roland Badeau, Bryan Pardo, Gaël Richard |
ICASSP | 4 |
| 2011 | A Computational Model of Auditory Perceptual Learning: Predicting Learning Interference Across Multiple Tasks
David Frank Little, Bryan Pardo, Beverly Wright |
CogSci | 2 |
| 2011 | A state space model for online polyphonic audio-score alignmentabstractWe present a novel online audio-score alignment approach for multi-instrument polyphonic music. This approach uses a 2-dimensional state vector to model the underlying score position and tempo of each time frame of the audio performance. The process model is defined by dynamic equations to transition between states. Two representations of the observed audio frame are proposed, resulting in two observation models: a multi-pitch-based and a chroma-based. Particle filtering is used to infer the hidden states from observations. Experiments on 150 music pieces with polyphony from one to four show the proposed approach outperforms an existing offline global string alignment-based score alignment approach. Results also show that the multi-pitch-based observation model works better than the chroma-based one. Zhiyao Duan, Bryan Pardo |
ICASSP | 2 |
| 2011 | Reconstructing completely overlapped notes from musical mixturesabstractIn mixtures of musical sounds, the problem of overlapped harmonics poses a significant challenge to source separation. Common Amplitude Modulation (CAM) is one of the most effective methods to resolve this problem. It, however, relies on non-overlapped harmonics from the same note being available. We propose an alternate technique for harmonic envelope estimation, based on Harmonic Temporal Envelope Similarity (HTES). We learn a harmonic envelope model for each instrument from the non-overlapped harmonics of notes of the same instrument, wherever they occur in the recording. This model is used to reconstruct the harmonic envelopes for overlapped harmonics. This allows reconstruction of completely overlapped notes. Experiments show our algorithm performs better than an existing system based on CAM when the harmonics of pitched instruments are strongly overlapped. Jinyu Han, Bryan Pardo |
ICASSP | 2 |
| 2011 | Degenerate Unmixing Estimation Technique using the Constant Q TransformabstractThe Degenerate Unmixing Estimation Technique (DUET) is a Blind Source Separation (BSS) algorithm for stereo audio. DUET depends on an amplitude-phase 2d histogram built from the differences between the two channels, where peaks in the histogram indicate sources in the mixture. If peaks overlap, separation becomes unfeasible. This is often the case for music mixtures. We propose to improve peak separation by building histograms from time-frequency representations based on the Constant Q Transform (CQT) instead of the Fourier Transform (FT). The CQT has a logarithmic frequency resolution matching the geometrically spaced notes of the Western music scale. We also adaptively resize histogram bins and use Wiener filtering to improve peak resolving and source reconstruction. Results on mixtures of harmonic musical instruments show improvement in separation, especially at low frequencies and for closely spaced sources. Zafar Rafii, Bryan Pardo |
ICASSP | 2 |
| 2011 | A simple music/voice separation method based on the extraction of the repeating musical structureabstractRepetition is a core principle in music. This is especially true for popular songs, generally marked by a noticeable repeating musical structure, over which the singer performs varying lyrics. On this basis, we propose a simple method for separating music and voice, by extraction of the repeating musical structure. First, the period of the repeating structure is found. Then, the spectrogram is segmented at period boundaries and the segments are averaged to create a repeating segment model. Finally, each time-frequency bin in a segment is compared to the model, and the mixture is partitioned using binary time-frequency masking by labeling bins similar to the model as the repeating background. Evaluation on a dataset of 1,000 song clips showed that this method can improve on the performance of an existing music/voice separation method without requiring particular features or complex frameworks. Zafar Rafii, Bryan Pardo |
ICASSP | 2 |
| 2010 | Song-level multi-pitch tracking by heavily constrained clusteringabstractGiven a set of monophonic, harmonic sound sources (e.g. human voices or wind instruments), multi-pitch estimation (MPE) is the task of determining the instantaneous pitches of each source. Multi-pitch tracking (MPT) connects the instantaneous pitch estimates provided by MPE algorithms into pitch trajectories of sources. A trajectory can be short (within a musical note), or long (an entire piece of music). While note-level MPT methods usually utilize local time-frequency proximity of pitches to connect them into a note, song-level MPT is much more difficult and needs more information. This is because pitches evolve discontinuously from note to note, and pitch trajectories can even interweave. In this paper, we cast the song-level MPT problem as a constrained clustering problem. The constraints are time-frequency locality of pitches and the clustering objective is their timbre consistency. Due to this problem's unique properties, existing constrained clustering algorithms cannot be directly applied. We propose a new constrained clustering algorithm. Experiments show that our approach produces good results on real-world music recordings of 4 musical instruments. Zhiyao Duan, Jinyu Han, Bryan Pardo |
ICASSP | 3 |
| 2010 | Multiple Fundamental Frequency Estimation by Modeling Spectral Peaks and Non-Peak RegionsabstractThis paper presents a maximum-likelihood approach to multiple fundamental frequency (F0) estimation for a mixture of harmonic sound sources, where the power spectrum of a time frame is the observation and the F0s are the parameters to be estimated. When defining the likelihood model, the proposed method models both spectral peaks and non-peak regions (frequencies further than a musical quarter tone from all observed peaks). It is shown that the peak likelihood and the non-peak region likelihood act as a complementary pair. The former helps find F0s that have harmonics that explain peaks, while the latter helps avoid F0s that have harmonics in non-peak regions. Parameters of these models are learned from monophonic and polyphonic training data. This paper proposes an iterative greedy search strategy to estimate F0s one by one, to avoid the combinatorial problem of concurrent F0 estimation. It also proposes a polyphony estimation method to terminate the iterative process. Finally, this paper proposes a postprocessing method to refine polyphony and F0 estimates using neighboring frames. This paper also analyzes the relative contributions of different components of the proposed method. It is shown that the refinement component eliminates many inconsistent estimation errors. Evaluations are done on ten recorded four-part J. S. Bach chorales. Results show that the proposed method shows superior F0 estimation and polyphony estimation compared to two state-of-the-art algorithms. Zhiyao Duan, Bryan Pardo, Changshui Zhang |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | 2DEQ: an intuitive audio equalizerabstractThe complexity of music production tools can be a significant bottleneck in the creative process. Here we describe the development of a simple, intuitive audio equalizer with the idea that our approach could also be applied to other types of music production tools. First, users generate a large set of equalization curves representative of the most common types of modifications. Next, we represent the entire set of curves in 2-dimensional space and determine the spatial location of common auditory adjectives. Finally we create an interface, called 2DEQ, where the user can drag a single dot to control equalization in this adjective-labeled space. Andrew T. Sabin, Bryan Pardo |
Creativity & Cognition | 2 |
| 2009 | Spectrum: Retrieving Different Points of View from the Blogosphere
Jiahui Liu 0002, Lawrence Birnbaum, Bryan Pardo |
ICWSM | 3 |
| 2009 | A method for rapid personalization of audio equalization parametersabstractPotential users of audio production software, such as audio equalizers, may be discouraged by the complexity of the interface. We describe a system that simplifies the interface by quickly mapping an individual's preferred sound manipulation onto parameters for audio equalization. This system learns mappings by presenting a sequence of equalizer settings to the user and correlating the gain in each frequency band with the user's preference rating. Learning typically converges in 25 user ratings (under two minutes). The system then creates a simple on-screen slider that lets the user manipulate the audio in terms of the descriptive term, without need to learn or use the parameters of an equalizer. Results are reported on the speed and effectiveness of the system for a set of 19 users and a set of five descriptive terms. Andrew T. Sabin, Bryan Pardo |
ACM Multimedia | 2 |
| 2009 | Classifying paintings by artistic genre: An analysis of features & classifiersabstractThis paper describes an approach to automatically classify digital pictures of paintings by artistic genre. While the task of artistic classification is often entrusted to human experts, recent advances in machine learning and multimedia feature extraction has made this task easier to automate. Automatic classification is useful for organizing large digital collections, for automatic artistic recommendation, and even for mobile capture and identification by consumers. Our evaluation uses variable resolution painting data gathered across Internet sources rather than solely using professional high-resolution data. Consequently, we believe this solution better addresses the task of classifying consumer-quality digital captures than other existing approaches. We include a comparison to existing feature extraction and classification methods as well as an analysis of our own approach across classifiers and feature vectors. Jana Zujovic, Lisa Gandy, Bryan Pardo, Thrasyvoulos N. Pappas |
MMSP | 4 |
| 2008 | Categorizing blogger's interests based on short snippets of blog postsabstractBlogs have become an important medium for people to express opinions and share information on the web. Predicting the interests of bloggers can be beneficial for information retrieval and knowledge discovery in the blogosphere. In this paper, we propose a two-layer classification model to categorize the interests of bloggers based on a set of short snippets collected from their blog posts. Experiments were conducted on a list of bloggers collected from blog directories, with their snippets collected from Google Blog Search. The results show that the proposed method is robust to errors in the lower level and achieve satisfactory performance in categorizing blogger's interests. Jiahui Liu 0002, Lawrence Birnbaum, Bryan Pardo |
CIKM | 3 |
| 2008 | Image spam hunterabstractSpammers are constantly creating sophisticated new weapons in their arms race with anti-spam technology, the latest of which is image-based spam. The newest image-based spam uses simple image processing technologies to vary the content of individual messages, e.g. by changing foreground colors, backgrounds, font types, or even rotating and adding artifacts to the images. Thus, they pose great challenges to conventional spam filters. In this paper, we propose a system using a probabilistic boosting tree to determine whether an incoming image is a spam or not based on global image features, i.e. color and gradient orientation histograms. The system identifies spam without the need for OCR and is robust in the face of the kinds of variation found in current spam images. Evaluation results show the system correctly classifies 90% of spam images while mislabeling only 0.86% of non-spam images as spam. Yan Gao 0003, Ming Yang 0007, Xiaonan Zhao, Bryan Pardo, Ying Wu 0001, Thrasyvoulos N. Pappas, Alok N. Choudhary |
ICASSP | 4 |
| 2007 | Towards a Model of Perceived Quality of Blind Audio Source SeparationabstractExisting perceptual models of audio quality, such as PEAQ, perform poorly when applied to blind audio source separation (BASS). We propose to create a perceptual model designed specifically for BASS algorithms. To create this model, we have designed a study to capture subjective human assessments of signal distortions resulting from BASS. In this study, humans rate the similarity between pairs of sounds. The first sound in each pair is a reference sound. The second sound is a distorted version of the reference, extracted from a multi-source mixture by a current BASS approach. We then correlate human similarity assessments with machine-measurable parameters. This paper describes preliminary results from a pilot study of three participants. Results indicate a strong correlation between human similarity assessments and the relative fraction of frames for which at last one frequency band in the distorted signal contains a significant noise component (RDF). Brendan Fox, Bryan Pardo |
ICME | 2 |
| 2007 | Learning to gesture: applying appropriate animations to spoken textabstractWe propose a machine learning system that learns to choose human gestures to accompany novel text. The system is trained on scripts comprised of speech and animations that were hand-coded by professional animators and shipped in video games. We treat this as a text-classification problem, classifying speech as corresponding with specific classes of gestures. We have built and tested two separate classifiers. The first is trained simply on the frequencies of different animations in the corpus. The second extracts text features from each script, and maps these features to the gestures that accompany the script. We have experimented with using a number of features of the text, including n-grams, emotional valence of the text, and parts-of-speech. Using a naïve Bayes classifier, the system learns to associate these features with appropriate classes of gestures. Once trained, the system can be given novel text for which it will attempt to assign appropriate gestures. We examine the performance of the two classifiers by using n-fold cross-validation over our training data, as well as two user studies of subjective evaluation of the results. Although there are many possible applications of automated gesture assignment, we hope to apply this technique to a system that produces an automated news show. Nathan D. Nichols, Jiahui Liu 0002, Bryan Pardo, Kristian J. Hammond, Lawrence Birnbaum |
ACM Multimedia | 3 |
| 2007 | A comparative evaluation of search techniques for query-by-humming using the MUSART testbedabstractAbstract Query‐by‐humming systems offer content‐based searching for melodies and require no special musical training or knowledge. Many such systems have been built, but there has not been much useful evaluation and comparison in the literature due to the lack of shared databases and queries. The MUSART project testbed allows various search algorithms to be compared using a shared framework that automatically runs experiments and summarizes results. Using this testbed, the authors compared algorithms based on string alignment, melodic contour matching, a hidden Markov model, n‐grams, and CubyHum. Retrieval performance is very sensitive to distance functions and the representation of pitch and rhythm, which raises questions about some previously published conclusions. Some algorithms are particularly sensitive to the quality of queries. Our queries, which are taken from human subjects in a realistic setting, are quite difficult, especially for n‐gram models. Finally, simulations on query‐by‐humming performance as a function of database size indicate that retrieval performance falls only slowly as the database size increases. Roger B. Dannenberg, William P. Birmingham, Bryan Pardo, Colin Meek, George Tzanetakis |
J. Assoc. Inf. Sci. Technol. | 3 |
| 2005 | Modeling Form for On-line Following of Musical Performances
Bryan Pardo, William P. Birmingham |
AAAI | 1 |
| 2005 | MusicStory: a personalized music video creatorabstractIn this paper, we describe MusicStory, a system that automatically creates videos to accompany music with lyrics. MusicStory uses common search engines, photo-sharing websites, and simple analysis of the dynamics and tempo of the music to create personalized photo-narratives. Video pacing and content is based on the content of the song and structure of the image repositories selected. The image associations MusicStory presents amplify the emotional experience by externalizing the imagery in song lyrics with the content found within a social network. The resulting work juxtaposes the meanings inherent in the social network with those in the song. David A. Shamma, Bryan Pardo, Kristian J. Hammond |
ACM Multimedia | 2 |
| 2004 | Name that tune: A pilot study in finding a melody from a sung queryabstractAbstract We have created a system for music search and retrieval. A user sings a theme from the desired piece of music. The sung theme (query) is converted into a sequence of pitch‐intervals and rhythms. This sequence is compared to musical themes (targets) stored in a database. The top pieces are returned to the user in order of similarity to the sung theme. We describe, in detail, two different approaches to measuring similarity between database themes and the sung query. In the first, queries are compared to database themes using standard string‐alignment algorithms. Here, similarity between target and query is determined by edit cost. In the second approach, pieces in the database are represented as hidden Markov models (HMMs). In this approach, the query is treated as an observation sequence and a target is judged similar to the query if its HMM has a high likelihood of generating the query. In this article we report our approach to the construction of a target database of themes, encoding, and transcription of user queries, and the results of preliminary experimentation with a set of sung queries. Our experiments show that while no approach is clearly superior to the other system, string matching has a slight advantage. Moreover, neither approach surpasses human performance. Bryan Pardo, Jonah Shifrin, William P. Birmingham |
J. Assoc. Inf. Sci. Technol. | 1 |
| 2003 | Changes in syllable magnitude and timing due to repeated correction
Caroline Menezes, Bryan Pardo, Donna Erickson, Osamu Fujimura |
Speech Commun. | 2 |