Paris Smaragdis

dblp:03/4048 · DBLP profile ↗
← Back
81ranked-venue papers
13as first author
19since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 61 · 5 first-author · 18 since 2021Artificial intelligence and machine learning · 26 · 7 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 On Class Separability Pitfalls In Audio-Text Contrastive Zero-Shot Learning
abstract
Recent advances in audio-text cross-modal contrastive learning have shown its potential towards zero-shot learning. One possibility for this is by projecting item embeddings from pre-trained backbone neural networks into a cross-modal space in which item similarity can be calculated in either domain. This process relies on a strong unimodal pre-training of the backbone networks, and on a data-intensive training task for the projectors. These two processes can be biased by unintentional data leakage, which can arise from using supervised learning in pre-training or from inadvertently training the cross-modal projection using labels from the zero-shot learning evaluation. In this study, we show that a significant part of the measured zero-shot learning accuracy is due to strengths inherited from the audio and text backbones, that is, they are not learned in the cross-modal domain and are not transferred from one modality to another.
Tiago Fernandes Tavares, Fábio J. Ayres, Zhepei Wang, Paris Smaragdis
ICASSP4
2024 Meta-AF Echo Cancellation for Improved Keyword Spotting
abstract
Adaptive filters (AFs) are vital for enhancing the performance of downstream tasks, such as speech recognition, sound event detection, and keyword spotting. However, traditional AF design prioritizes isolated signal-level objectives, often overlooking downstream task performance. This can lead to suboptimal performance. Recent research has leveraged meta-learning to automatically learn AF update rules from data, alleviating the need for manual tuning when using simple signal-level objectives. This paper improves the Meta-AF [1] framework by expanding it to support end-to-end training for arbitrary downstream tasks. We focus on classification tasks, where we introduce a novel training methodology that harnesses self-supervision and classifier feedback. We evaluate our approach on the combined task of acoustic echo cancellation and keyword spotting. Our findings demonstrate consistent performance improvements with both pre-trained and joint-trained keyword spotting models across synthetic and real playback. Notably, these improvements come without requiring additional tuning, increased inference-time complexity, or reliance on oracle signal-level training data.
Jonah Casebeer, Junkai Wu, Paris Smaragdis
ICASSP3
2024 Noise-Robust DSP-Assisted Neural Pitch Estimation With Very Low Complexity
abstract
Pitch estimation is an essential step of many speech processing algorithms, including speech coding, synthesis, and enhancement. Recently, pitch estimators based on deep neural networks (DNNs) have been outperforming well-established DSP-based techniques. Unfortunately, these new estimators can be impractical to deploy in real-time systems, both because of their relatively high complexity, and the fact that some require significant lookahead. We show that a hybrid estimator using a small deep neural network (DNN) with traditional DSP-based features can match or exceed the performance of pure DNN-based models, with a complexity and algorithmic delay comparable to traditional DSP-based algorithms. We further demonstrate that this hybrid approach can provide benefits for a neural vocoding task.
Krishna Subramani, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP4
2024 Audio Editing with Non-Rigid Text Prompts
Francesco Paissan, Luca Della Libera, Zhepei Wang, Paris Smaragdis, Mirco Ravanelli, Cem Subakan
INTERSPEECH4
2023 Latent Iterative Refinement for Modular Source Separation
abstract
Traditional source separation approaches train deep neural network models end-to-end with all the data available at once by minimizing the empirical risk on the whole training set. On the inference side, after training the model, the user fetches a static computation graph and runs the full model on some specified observed mixture signal to get the estimated source signals. Additionally, many of those models consist of several basic processing blocks which are applied sequentially. We argue that we can significantly increase resource efficiency during both training and inference stages by re-formulating a model’s training and inference procedures as iterative mappings of latent signal representations. First, we can apply the same processing block more than once on its output to refine the input signal and consequently improve parameter efficiency. During training, we can follow a block-wise procedure which enables a reduction on memory requirements. Thus, one can train a very complicated network structure using significantly less computation compared to end-to-end training. During inference, we can dynamically adjust how many processing blocks and iterations of a specific block an input signal needs using a gating module.
Dimitrios Bralios, Efthymios Tzinis, Gordon Wichern, Paris Smaragdis, Jonathan Le Roux
ICASSP4
2023 Generative Modeling Based Manifold Learning for Adaptive Filtering Guidance
abstract
In most practical adaptive filtering problems, estimated filters are not arbitrary, but instead lie on a manifold that encapsulates characteristics of the problem at hand. Consequently, it is desirable to steer adaptation towards filters that lie on that manifold. In this paper, we propose a novel approach to learn the manifold of a set of impulse responses and subsequently employ that learned manifold in an adaptation algorithm for system identification. The presented approach is a practical adaptive filtering recipe for enforcing a data-driven search domain constraint, instead of using conventional constrained optimization methods.
Karim Helwani, Paris Smaragdis, Michael M. Goodwin
ICASSP2
2023 Framewise Wavegan: High Speed Adversarial Vocoder In Time Domain With Very Low Computational Complexity
abstract
GAN vocoders are currently one of the state-of-the-art methods for building high-quality neural waveform generative models. However, most of their architectures require dozens of billion floating-point operations per second (GFLOPS) to generate speech waveforms in samplewise manner. This makes GAN vocoders still challenging to run on normal CPUs without accelerators or parallel computers. In this work, we propose a new architecture for GAN vocoders that mainly depends on recurrent and fully-connected networks to directly generate the time domain signal in framewise manner. This results in considerable reduction of the computational cost and enables very fast generation on both GPUs and low-complexity CPUs. Experimental results show that our Framewise WaveGAN vocoder achieves significantly higher quality than auto-regressive maximum-likelihood vocoders such as LPCNet at a very low complexity of 1.2GFLOPS. This makes GAN vocoders more practical on edge and low-power devices.
Ahmed Mustafa, Jean-Marc Valin, Jan Büthe, Paris Smaragdis, Michael M. Goodwin
ICASSP4
2023 Optimal Condition Training for Target Source Separation
abstract
Recent research has shown remarkable performance in leveraging multiple extraneous conditional and non-mutually-exclusive semantic concepts for sound source separation, allowing the flexibility to extract a given target source based on multiple different queries. In this work, we propose a new optimal condition training (OCT) method for single-channel target source separation, based on greedy parameter updates using the highest performing condition among equivalent conditions associated with a given target source. Our experiments show that the complementary information carried by the diverse semantic concepts significantly helps to disentangle and isolate sources of interest much more efficiently compared to single-conditioned models. Moreover, we propose a variation of OCT with condition refinement, in which an initial condition vector is adapted to the given mixture and transformed to a more amenable representation for target source extraction. We showcase the effectiveness of OCT on diverse source separation experiments where it improves upon permutation invariant models with oracle assignment between estimated and target sources and obtains state-of-the-art performance in the more challenging task of text-based source separation, outperforming even dedicated text-only conditioned models.
Efthymios Tzinis, Gordon Wichern, Paris Smaragdis, Jonathan Le Roux
ICASSP3
2023 A Framework for Unified Real-Time Personalized and Non-Personalized Speech Enhancement
abstract
In this study, we present an approach to train a single speech enhancement network that can perform both personalized and non-personalized speech enhancement. This is achieved by incorporating a frame-wise conditioning input that specifies the type of enhancement output. To improve the quality of the enhanced output and mitigate oversuppression, we experiment with re-weighting frames by the presence or absence of speech activity and applying augmentations to speaker embeddings. By training under a multi-task learning setting, we empirically show that the proposed unified model obtains promising results on both personalized and non-personalized speech enhancement benchmarks and reaches similar performance to models that are trained specialized for either task. The strong performance of the proposed method demonstrates that the unified model is a more economical alternative compared to keeping separate task-specific models during inference.
Zhepei Wang, Ritwik Giri, Devansh Shah, Jean-Marc Valin, Michael M. Goodwin, Paris Smaragdis
ICASSP6
2023 Meta-AF: Meta-Learning for Adaptive Filters
abstract
Adaptive filtering algorithms are pervasive throughout signal processing and have had a material impact on a wide variety of domains including audio processing, telecommunications, biomedical sensing, astrophysics and cosmology, seismology, and many more. Adaptive filters typically operate via specialized online, iterative optimization methods such as least-mean squares or recursive least squares and aim to process signals in unknown or nonstationary environments. Such algorithms, however, can be slow and laborious to develop, require domain expertise to create, and necessitate mathematical insight for improvement. In this work, we seek to improve upon hand-derived adaptive filter algorithms and present a comprehensive framework for learning online, adaptive signal processing algorithms or update rules directly from data. To do so, we frame the development of adaptive filters as a meta-learning problem in the context of deep learning and use a form of self-supervision to learn online iterative update rules for adaptive filters. To demonstrate our approach, we focus on audio applications and systematically develop meta-learned adaptive filters for five canonical audio problems including system identification, acoustic echo cancellation, blind equalization, multi-channel dereverberation, and beamforming.We compare our approach against common baselines and/or recent state-of-the-art methods. We show we can learn high-performing adaptive filters that operate in real-time and, in most cases, significantly outperform each method we compare against – all using a single general-purpose configuration of our approach.
Jonah Casebeer, Nicholas J. Bryan, Paris Smaragdis
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Neural Speech Synthesis on a Shoestring: Improving the Efficiency of Lpcnet
abstract
Neural speech synthesis models can synthesize high quality speech but typically require a high computational complexity to do so. In previous work, we introduced LPCNet, which uses linear prediction to significantly reduce the complexity of neural synthesis. In this work, we further improve the efficiency of LPCNet – targeting both algorithmic and computational improvements – to make it usable on a wide variety of devices. We demonstrate an improvement in synthesis quality while operating 2.5x faster. The resulting open-source1LPCNet algorithm can perform real-time neural synthesis on most existing phones and is even usable in some embedded devices.
Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy
ICASSP3
2022 End-to-end LPCNet: A Neural Vocoder With Fully-Differentiable LPC Estimation
abstract
Neural vocoders have recently demonstrated high quality speech synthesis, but typically require a high computational complexity.LPCNet was proposed as a way to reduce the complexity of neural synthesis by using linear prediction (LP) to assist an autoregressive model.At inference time, LPCNet relies on the LP coefficients being explicitly computed from the input acoustic features.That makes the design of LPCNet-based systems more complicated, while adding the constraint that the input features must represent a clean speech spectrum.We propose an end-to-end version of LPCNet that lifts these limitations by learning to infer the LP coefficients from the input features in the frame rate network .Results show that the proposed end-toend approach equals or exceeds the quality of the original LPC-Net model, but without explicit LP analysis.Our open-source 1 end-to-end model still benefits from LPCNet's low complexity, while allowing for any type of conditioning features.
Krishna Subramani, Jean-Marc Valin, Umut Isik, Paris Smaragdis, Arvindh Krishnaswamy
INTERSPEECH4
2022 Heterogeneous Target Speech Separation
abstract
We introduce a new paradigm for single-channel target source separation where the sources of interest can be distinguished using non-mutually exclusive concepts (e.g., loudness, gender, language, spatial location, etc).Our proposed heterogeneous separation framework can seamlessly leverage datasets with large distribution shifts and learn cross-domain representations under a variety of concepts used as conditioning.Our experiments show that training separation models with heterogeneous conditions facilitates the generalization to new concepts with unseen out-of-domain data while also performing substantially higher than single-domain specialist models.Notably, such training leads to more robust learning of new harder source separation discriminative concepts and can yield improvements over permutation invariant training with oracle source selection.We analyze the intrinsic behavior of source separation training with heterogeneous metadata and propose ways to alleviate emerging problems with challenging separation conditions.We release the collection of preparation recipes for all datasets used to further promote research towards this challenging task.
Efthymios Tzinis, Gordon Wichern, Aswin Shanmugam Subramanian, Paris Smaragdis, Jonathan Le Roux
INTERSPEECH4
2022 Real-Time Packet Loss Concealment With Mixed Generative and Predictive Model
Jean-Marc Valin, Ahmed Mustafa, Christopher Montgomery, Timothy B. Terriberry, Michael Klingbeil, Paris Smaragdis, Arvindh Krishnaswamy
INTERSPEECH6
2022 Learning Representations for New Sound Classes With Continual Self-Supervised Learning
abstract
In this article, we work on a sound recognition system that continually incorporates new sound classes. Our main goal is to develop a framework where the model can be updated without relying on labeled data. For this purpose, we propose adopting representation learning, where an encoder is trained using unlabeled data. This learning framework enables the study and implementation of a practically relevant use case where only a small amount of the labels is available in a continual learning context. We also make the empirical observation that a similarity-based representation learning method within this framework is robust to forgetting even if no explicit mechanism against forgetting is employed. We show that this approach obtains similar performance compared to several distillation-based continual learning methods when employed on self-supervised representation learning methods.
Zhepei Wang, Cem Subakan, Xilin Jiang, Junkai Wu, Efthymios Tzinis, Mirco Ravanelli, Paris Smaragdis
IEEE Signal Process. Lett.7
2021 Communication-Cost Aware Microphone Selection for Neural Speech Enhancement with Ad-Hoc Microphone Arrays
abstract
In this paper, we present a method for jointly-learning a microphone selection mechanism and a speech enhancement network for multi-channel speech enhancement with an ad-hoc microphone array. The attention-based microphone selection mechanism is trained to reduce communication-costs through a penalty term which represents a task-performance/ communication-cost trade-off. While working within the trade-off, our method can intelligently stream from more microphones in lower SNR scenes and fewer microphones in higher SNR scenes. We evaluate the model in complex echoic acoustic scenes with moving sources and show that it matches the performance of models that stream from a fixed number of microphones while reducing communication costs.
Jonah Casebeer, Jamshed Kaikaus, Paris Smaragdis
ICASSP3
2021 Differentiable Signal Processing With Black-Box Audio Effects
abstract
We present a data-driven approach to automate audio signal processing by incorporating stateful third-party, audio effects as layers within a deep neural network. We then train a deep encoder to analyze input audio and control effect parameters to perform the desired signal manipulation, requiring only input-target paired audio data as supervision. To train our network with non-differentiable black-box effects layers, we use a fast, parallel stochastic gradient approximation scheme within a standard auto differentiation graph, yielding efficient end-to-end backpropagation. We demonstrate the power of our approach with three separate automatic audio production applications: tube amplifier emulation, automatic removal of breaths and pops from voice recordings, and automatic music mastering. We validate our results with a subjective listening test, showing our approach not only can enable new automatic audio effects tasks, but can yield results comparable to a specialized, state-of-the-art commercial solution for music mastering.
Marco A. Martínez Ramírez, Oliver Wang, Paris Smaragdis, Nicholas J. Bryan
ICASSP3
2021 Unified Gradient Reweighting for Model Biasing with Applications to Source Separation
abstract
Recent deep learning approaches have shown great improvement in audio source separation tasks. However, the vast majority of such work is focused on improving average separation performance, often neglecting to examine or control the distribution of the results. In this paper, we propose a simple, unified gradient reweighting scheme, with a lightweight modification to bias the learning process of a model and steer it towards a certain distribution of results. More specifically, we reweight the gradient updates of each batch, using a user-specified probability distribution. We apply this method to various source separation tasks, in order to shift the operating point of the models towards different objectives. We demonstrate different parameterizations of our unified reweighting scheme can be used towards addressing several real-world problems, such as unreliable separation estimates. Our framework enables the user to control a robustness trade-off between worst and average performance. Moreover, we experimentally show that our unified reweighting scheme can also be used in order to shift the focus of the model towards being more accurate for user-specified sound classes or even towards easier examples in order to enable faster convergence.
Efthymios Tzinis, Dimitrios Bralios, Paris Smaragdis
ICASSP3
2021 Optimizing Short-Time Fourier Transform Parameters via Gradient Descent
abstract
The Short-Time Fourier Transform (STFT) has been a staple of signal processing, often being the first step for many audio tasks. A very familiar process when using the STFT is the search for the best STFT parameters, as they often have significant side effects if chosen poorly. These parameters are often defined in terms of an integer number of samples, which makes their optimization non-trivial. In this paper we show an approach that allows us to obtain a gradient for STFT parameters with respect to arbitrary cost functions, and thus enable the ability to employ gradient descent optimization of quantities like the STFT window length, or the STFT hop size. We do so for parameter values that stay constant throughout an input, but also for cases where these parameters have to dynamically change over time to accommodate varying signal characteristics.
Krishna Subramani, Paris Smaragdis
ICASSP3
2020 One-Shot Parametric Audio Production Style Transfer with Application to Frequency Equalization
abstract
Audio production is a difficult process for many people], [and properly manipulating sound to achieve a certain effect is non-trivial. In this paper], [we present a method that facilitates this process by inferring appropriate audio effect parameters in order to make an input recording sound similar to an unrelated reference recording. We frame our work as a form of parametric style transfer that], [by design], [leverages existing audio production semantics and manipulation algorithms], [avoiding several issues that have plagued audio style transfer algorithms in the past. To demonstrate our approach], [we consider the task of controlling a parametric], [four-band infinite impulse response equalizer and show that we are able to predict the parameters necessary to transform the equalization style of one recording to another. The framework we present], [however], [is applicable to a wider range of parametric audio effects.
Stylianos I. Mimilakis, Nicholas J. Bryan, Paris Smaragdis
ICASSP3
2020 Two-Step Sound Source Separation: Training On Learned Latent Targets
abstract
In this paper, we propose a two-step training procedure for source separation via a deep neural network. In the first step we learn a transform (and it's inverse) to a latent space where masking-based separation performance using oracles is optimal. For the second step, we train a separation module that operates on the previously learned space. In order to do so, we also make use of a scale-invariant signal to distortion ratio (SI-SDR) loss function that works in the latent space, and we prove that it lower-bounds the SI-SDR in the time domain. We run various sound separation experiments that show how this approach can obtain better performance as compared to systems that learn the transform and the separation module jointly. The proposed methodology is general enough to be applicable to a large class of neural network end-to-end separation systems.
Efthymios Tzinis, Shrikant Venkataramani, Zhepei Wang, Y. Cem Sübakan, Paris Smaragdis
ICASSP5
2020 End-To-End Non-Negative Autoencoders for Sound Source Separation
abstract
Discriminative models for source separation have recently been shown to produce impressive results. However, when operating on sources outside of the training set, these models can not perform as well and are cumbersome to update. Classical methods like Nonnegative Matrix Factorization (NMF) provide modular approaches to source separation that can be easily updated to adapt to new mixture scenarios. In this paper, we generalize NMF to develop end-to-end non-negative auto-encoders and demonstrate how they can be used for source separation. Our experiments indicate that these models deliver comparable separation performance to discriminative approaches, while retaining the modularity of NMF and the modeling flexibility of neural networks.
Shrikant Venkataramani, Efthymios Tzinis, Paris Smaragdis
ICASSP3
2019 VoiceAssist: Guiding Users to High-Quality Voice Recordings
abstract
Voice recording is a challenging task with many pitfalls due to sub-par recording environments, mistakes in recording setup, microphone quality, etc. Newcomers to voice recording often have difficulty recording their voice, leading to recordings with low sound quality. Many amateur recordings of poor quality have two key problems: too much reverberation (echo), and too much background noise (e.g. fans, electronics, street noise). We present VoiceAssist, a system that helps inexperienced users produce high quality recordings by providing real-time visual feedback on audio quality. We integrate modern audio quality measures into an interactive human-machine feedback loop, so that the audio quality can be maximized at capture-time. We demonstrate the utility of this feedback for improving the recording quality with a user study. When presented with visual feedback about recording quality, users produced recordings that were strongly preferred by third-party listeners, when compared to recordings made without this feedback.
Prem Seetharaman, Gautham J. Mysore, Bryan Pardo, Paris Smaragdis, Celso Gomes
CHI4
2019 Multi-view Networks for Multi-channel Audio Classification
abstract
In this paper we introduce the idea of multi-view networks for sound classification with multiple sensors. We show how one can build a multi-channel sound recognition model trained on a fixed number of channels, and deploy it to scenarios with arbitrary (and potentially dynamically changing) number of input channels and not observe degradation in performance. We demonstrate that at inference time you can safely provide this model all available channels as it can ignore noisy information and leverage new information better than standard baseline approaches. The model is evaluated in both an anechoic environment and in rooms generated by a room acoustics simulator. We demonstrate that this model can generalize to unseen numbers of channels as well as unseen room geometries.
Jonah Casebeer, Zhepei Wang, Paris Smaragdis
ICASSP3
2019 Majorization-minimization Algorithms for Convolutive NMF with the Beta-divergence
abstract
Nonnegative matrix factorization (NMF) has become a method of choice for spectrogram decomposition. However, its inability to capture dependencies across columns of the input motivated the introduction of a variant, convolutive NMF. While algorithms for solving the convolutive NMF problem were previously proposed, they rely on the use of a heuristic that does not insure the convergence of the algorithm (in particular in terms of objective function values). The goal of this work is to propose rigorous update rules, based on a majorization-minimization (MM) approach, for convolutive NMF with the β-divergence (a standard family of measures of fit). Specifically, we derive and study two variants of a convolutive NMF algorithm that are guaranteed to decrease the objective function value at each iteration. The complexity of the algorithms is studied, and the performance in terms of execution time and objective function are evaluated and compared in several numerical experiments using real-world audio data. Experiments show that the proposed MM algorithms consistently provide lower values of the objective function than the heuristic, at similar computational cost.
Dylan Fagot, Herwig Wendt, Cédric Févotte, Paris Smaragdis
ICASSP4
2019 Unsupervised Deep Clustering for Source Separation: Direct Learning from Mixtures Using Spatial Information
abstract
We present a monophonic source separation system that is trained by only observing mixtures with no ground truth separation information. We use a deep clustering approach which trains on multichannel mixtures and learns to project spectrogram bins to source clusters that correlate with various spatial features. We show that using such a training process we can obtain separation performance that is as good as making use of ground truth separation information. Once trained, this system is capable of performing sound separation on monophonic inputs, despite having learned how to do so using multi-channel recordings.
Efthymios Tzinis, Shrikant Venkataramani, Paris Smaragdis
ICASSP3
2018 Bitwise Neural Networks for Efficient Single-Channel Source Separation
abstract
We present Bitwise Neural Networks (BNN) as an efficient hardware-friendly solution to single-channel source separation tasks in resource-constrained environments. In the proposed BNN system, we replace all the real-valued operations during the feedforward process of a Deep Neural Network (DNN) with bitwise arithmetic (e.g. the XNOR operation between bipolar binaries in place of multiplications). Thanks to the fully bitwise run-time operations, the BNN system can serve as an alternative solution where efficient real-time processing is critical, for example real-time speech enhancement in embedded systems. Furthermore, we also propose a binarization scheme to convert the input signals into bit strings so that the BNN parameters learn the Boolean mapping between input binarized mixture signals and their target Ideal Binary Masks (IBM). Experiments on the single-channel speech denoising tasks show that the efficient BNN-based source separation system works well with an acceptable performance loss compared to a comprehensive real-valued network, while consuming a minimal amount of resources.
Minje Kim 0001, Paris Smaragdis
ICASSP2
2018 Blind Estimation of the Speech Transmission Index for Speech Quality Prediction
abstract
The speech transmission index (STI) of a listening position within a given room indicates the quality and intelligibility of speech uttered in that room. The measure is very reliable for predicting speech intelligibility in many room conditions but requires an STI measurement of the impulse response for the room. We present a method for blindly estimating the STI without measuring or modeling the impulse response of the room using deep convolutional neural networks. Our model is trained entirely using simulated room impulse responses combined with clean speech examples from the DAPS dataset [1] and works directly on PCM audio. Our experiments show that our method predicts true STI with a high degree of accuracy - an average error of under 4%. It can also distinguish between different STI conditions to a level of granularity that is comparable to humans.
Prem Seetharaman, Gautham J. Mysore, Paris Smaragdis, Bryan Pardo
ICASSP3
2018 Generative Adversarial Source Separation
abstract
Generative source separation methods such as non-negative matrix factorization (NMF) or auto-encoders, rely on the assumption of an output probability density. Generative Adversarial Networks (GANs) can learn data distributions without needing a parametric assumption on the output density. We show on a speech source separation experiment that, a multilayer perceptron trained with a Wasserstein-GAN formulation outperforms NMF, auto-encoders trained with maximum likelihood, and variational auto-encoders in terms of source to distortion ratio.
Y. Cem Sübakan, Paris Smaragdis
ICASSP2
2017 A neural network alternative to non-negative audio models
abstract
We present a neural network that can act as an equivalent to a Non-Negative Matrix Factorization (NMF), and further show how it can be used to perform supervised source separation. Due to the extensibility of this approach we show how we can achieve better source separation performance as compared to NMF-based methods, and propose a variety of derivative architectures that can be used for further improvements.
Paris Smaragdis, Shrikant Venkataramani
ICASSP1
2017 AutoDub: Automatic Redubbing for Voiceover Editing
abstract
Redubbing is an extensively used technique to correct errors in voiceover recordings. It involves re-recording a part of a voiceover, identifying the corresponding section of audio in the original recording that needs to be replaced, and using low level audio tools to replace the audio. Although this sequence of steps can be performed using traditional audio editing tools, the process can be tedious when dealing with long voiceover recordings and prohibitively difficult for users not familiar with such tools. To address this issue, we present AutoDub, a novel system for redubbing voiceover recordings. Using our system, a user simply needs to re-record the part of the voiceover that needs to be replaced. Our system automatically locates the corresponding part in the original recording and performs the low level audio processing to replace it. The system can be easily incorporated in any existing sophisticated audio editor or can be employed as a functionality in an audio-guided user interface. User studies involving participation from novice, knowledgeable and expert users indicate that our tool is preferred to a traditional audio editor based redubbing approach by all categories of users due to its faster and easier redubbing capabilities.
Shrikant Venkataramani, Paris Smaragdis, Gautham J. Mysore
UIST2
2016 Efficient neighborhood-based topic modeling for collaborative audio enhancement on massive crowdsourced recordings
abstract
Collaborative Audio Enhancement (CAE) aims at separating a dominant source from crowdsourced recordings of a scene. This paper proposes a CAE setup as a big ad-hoc microphone array problem, assuming hundreds of sensors scattered over a large scene, e.g. a concert hall or a street riot. An important characteristic in such cases is the fact that not all sensors capture useful information, mainly because of the existence of strong local noise interferences and recording artifacts. This renders traditional array processing techniques inadequate for tasks such as source enhancement. One way to recover the most common source while suppressing recording-specific interference, is to share latent components across simultaneous models on multiple magnitude spectrograms. The proposed method improves on the quality and the computational requirements of such a model by using a two-stage nearest-neighborhood search at every EM update. Its optional first-round search uses Hamming distance between hashed spectrograms to quickly find a redundant candidate set, and then a subsequent step narrows the set down to a subset using more appropriate cross entropy. Experimental results show that the proposed neighborhood schemes converge to the better quality solutions faster than the comprehensive model using all data.
Minje Kim 0001, Paris Smaragdis
ICASSP2
2016 Robust Source Localization and Enhancement With a Probabilistic Steered Response Power Model
abstract
Source localization and enhancement are often treated separately in the array processing literature. One can apply steered response power (SRP) localization to determine the sources' Directions-Of-Arrival (DOA) followed by beamforming and Wiener post-filtering to isolate the sources from each other and ambient interference. We show that when there is significant overlap between directional sources of interest in the time-frequency (TF) plane, traditional SRP localization breaks down. This may occur, for example, when the array is located near a reflector, significant early reflections are present, or the sources are harmonized. We propose a joint solution to the localization and enhancement problems via a probabilistic interpretation of the SRP function. We formulate optimization procedures for (1) a mixture of single-source SRP distributions (MoSRP) and (2) a multi-source SRP distribution (MultSRP). Unlike in traditional localization, the latter approach explicitly models source overlap in the TF plane. Results shows that the MultSRP model is capable of localizing sources with significant overlap in the TF domain and that either of the proposed methods out-performs standard SRP localization for multiple speakers.
Johannes Traa, David Wingate, Noah D. Stein, Paris Smaragdis
IEEE ACM Trans. Audio Speech Lang. Process.4
2015 Efficient manifold preserving audio source separation using locality sensitive hashing
abstract
We propose an efficient technique to learn probabilistic hierarchical topic models that are designed to preserve the manifold structure of audio data. The consideration of the data manifold is important, as it has been shown to provide superior performance in certain audio applications such as source separation. However, the high computational cost of a sparse encoding step due to the requirement of a large dictionary prevents it from being used in real-world applications such as real-time speech enhancement and the analysis of big audio data. In order to achieve a substantial speed-up of this step, while still respecting the data manifold, we propose to harmonize a particular type of locality sensitive hashing with the hierarchical topic model. The proposed use of hashing can reduce the computational complexity of the sparse encoding by providing candidates of non-zero activations, where the candidate set is built based on Hamming distance. The hashing step is followed by comprehensive sparse coding that considers those candidates only, rather than the entire dictionary. Experimental results show that the proposed hashing technique can provide audio source separation results comparable to the similar system without hashing, but with significantly less and cheaper computation.
Minje Kim 0001, Paris Smaragdis, Gautham J. Mysore
ICASSP2
2015 Joint acoustic and spectral modeling for speech dereverberation using non-negative representations
abstract
This paper proposes a single-channel speech dereverberation method enhancing the spectrum of the reverberant speech signal. The proposed method uses a non-negative approximation of the convolutive transfer function (N-CTF) to simultaneously estimate the magnitude spectrograms of the speech signal and the room impulse response (RIR). To utilize the speech spectral structure, we propose to model the speech spectrum using non-negative matrix factorization, which is directly used in the N-CTF model resulting in a new cost function. We derive new estimators for the parameters by minimizing the obtained cost function. Additionally, to investigate the effect of the speech temporal dynamics for dereverberation, we use a frame stacking method and derive optimal estimators. Experiments are performed for two measured RIRs and the performance of the proposed method is compared to the performance of a state-of-the-art dereverberation method enhancing the speech spectrum. Experimental results show that the proposed method improved instrumental speech quality measures, where using speech temporal dynamics was found to be beneficial in severe reverberation conditions.
Nasser Mohammadiha, Paris Smaragdis, Simon Doclo
ICASSP2
2015 Mixtures of Local Dictionaries for Unsupervised Speech Enhancement
abstract
We propose a novel extension of Nonnegative Matrix Factorization (NMF) that models a signal with multiple local dictionaries activated sparsely. This set of local dictionaries for a source, e.g., speech, disjointly constitute a superset that is more discriminative than an ordinary NMF dictionary, because its local structures represent the source's manifold better. A block sparsity constraint is used to regularize the NMF solutions so that only one or a small number of blocks are active at a given time. Moreover, a concentrationz prior further regularizes each block of bases to be close to each other for better locality preservation. We test the proposed Mixture of Local Dictionaries (MLD) on single-channel speech enhancement tasks and show that it outperforms the state of the art technology by up to 2 dB in signal-to-distortion ratio, especially in the unsupervised environment where neither the speaker identity nor the type of noise is known in advance.
Minje Kim 0001, Paris Smaragdis
IEEE Signal Process. Lett.2
2015 Joint Optimization of Masks and Deep Recurrent Neural Networks for Monaural Source Separation
abstract
Monaural source separation is important for many real world applications. It is challenging because, with only a single channel of information available, without any constraints, an infinite number of solutions are possible. In this paper, we explore joint optimization of masking functions and deep recurrent neural networks for monaural source separation tasks, including speech separation, singing voice separation, and speech denoising. The joint optimization of the deep recurrent neural networks with an extra masking layer enforces a reconstruction constraint. Moreover, we explore a discriminative criterion for training neural networks to further enhance the separation performance. We evaluate the proposed system on the TSP, MIR-1K, and TIMIT datasets for speech separation, singing voice separation, and speech denoising tasks, respectively. Our approaches achieve 2.30-4.98 dB SDR gain compared to NMF models in the speech separation task, 2.30-2.48 dB GNSDR gain and 4.32-5.42 dB GSIR gain compared to existing models in the singing voice separation task, and outperform NMF and DNN baselines in the speech denoising task.
Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Deep learning for monaural speech separation
abstract
Monaural source separation is useful for many real-world applications though it is a challenging problem. In this paper, we study deep learning for monaural speech separation. We propose the joint optimization of the deep learning models (deep neural networks and recurrent neural networks) with an extra masking layer, which enforces a reconstruction constraint. Moreover, we explore a discriminative training criterion for the neural networks to further enhance the separation performance. We evaluate our approaches using the TIMIT speech corpus for a monaural speech separation task. Our proposed models achieve about 3.8∼4.9 dB SIR gain compared to NMF models, while maintaining better SDRs and SARs.
Po-Sen Huang, Minje Kim 0001, Mark Hasegawa-Johnson, Paris Smaragdis
ICASSP4
2014 Phase and level difference fusion for robust multichannel source separation
abstract
Inter-channel phase (IPD) and level (ILD) differences are common features in multichannel source separation algorithms like DUET and MENUET. However, their utility depends strongly on the configuration of the array and what microphone pairs are used to calculate them. IPDs are most useful when extracted from microphones that are close together as this avoids spatial aliasing. In contrast, ILD clusters are only well separated for widely spaced microphones. We investigate this trade-off between IPD and ILD features and propose a method to best combine them for multichannel source separation. Experimental results demonstrate the utility of this approach.
Johannes Traa, Minje Kim 0001, Paris Smaragdis
ICASSP3
2014 Experiments on deep learning for speech denoising
abstract
In this paper we present some experiments using a deep learn-ing model for speech denoising. We propose a very lightweight procedure that can predict clean speech spectra when presented with noisy speech inputs, and we show how various parameter choices impact the quality of the denoised signal. Through our experiments we conclude that such a structure can perform bet-ter than some comparable single-channel approaches and that it is able to generalize well across various speakers, noise types and signal-to-noise ratios.
Paris Smaragdis, Minje Kim 0001
INTERSPEECH2
2014 Spectral Learning of Mixture of Hidden Markov Models
Y. Cem Sübakan, Johannes Traa, Paris Smaragdis
NIPS3
2014 Multichannel source separation and tracking with RANSAC and directional statistics
abstract
We describe multichannel blind source separation and tracking algorithms based on clustering wrapped interchannel phase difference (IPD) features. We pose the clustering problem as one of multimodal circular-linear regression and present its probabilistic formulation. Phase wrapping due to spatial aliasing is explicitly incorporated by modeling the IPD features as circular variables. We present two methods based on Expectation-Maximization (EM) and a sequential variant of RANdom SAmple Consensus (RANSAC). We show that their strengths can be combined by using RANSAC to initialize EM. The IPD clustering algorithm is applied to separate stationary speakers from a multichannel mixture. We then extend it to the case of moving speakers by tracking their directions-of-arrival with the Factorial Wrapped Kalman Filter (FWKF) using RANSAC as a data preprocessor. Experimental results demonstrate that the proposed methods perform well in the presence of reverberant babble noise and spatial aliasing. The FWKF successfully tracks and separates moving speakers with separation quality comparable to that for stationary speakers.
Johannes Traa, Paris Smaragdis
IEEE ACM Trans. Audio Speech Lang. Process.2
2013 Collaborative audio enhancement using probabilistic latent component sharing
abstract
This paper presents a collaborative audio enhancement system that aims to recover common audio sources from multiple recordings of a given audio scene. We do so in the context where each recording is uniquely corrupted. To this end, we propose a method of simultaneous probabilistic latent component analyses on synchronized inputs. In the proposed model, some of the parameters are fixed to be same during and after the learning process to capture common audio content while the rest models unwanted recording-specific interferences and artifacts. Our model also allows for prior knowledge about the parameters of the model, e.g. representative spectra of the components, to be incorporated in the factorization. A post processing scheme that consolidates the extracted sources from the set of inputs is also proposed to handle the possible loss of certain frequency regions. Experiments on commercial music signals with various artifacts show the merit of the proposed method.
Minje Kim 0001, Paris Smaragdis
ICASSP2
2013 Prediction based filtering and smoothing to exploit temporal dependencies in NMF
abstract
Nonnegative matrix factorization is an appealing technique for many audio applications. However, in it's basic form it does not use temporal structure, which is an important source of information in speech processing. In this paper, we propose NMF-based filtering and smoothing algorithms that are related to Kalman filtering and smoothing. While our prediction step is similar to that of Kalman filtering, we develop a multiplicative update step which is more convenient for nonnegative data analysis and in line with existing NMF literature. The proposed smoothing approach introduces an unavoidable processing delay, but the filtering algorithm does not and can be readily used for on-line applications. Our experiments using the proposed algorithms show a significant improvement over the baseline NMF approaches. In the case of speech denoising with factory noise at 0 dB input SNR, the smoothing algorithm outperforms NMF with 3.2 dB in SDR and around 0.5 MOS in PESQ, likewise source separation experiments result in improved performance due to taking advantage of the temporal regularities in speech.
Nasser Mohammadiha, Paris Smaragdis, Arne Leijon
ICASSP2
2013 Blind multi-channel source separation by circular-linear statistical modeling of phase differences
abstract
We address the problem of blind separation of speech signals with a microphone array. We demonstrate that a signal propagating towards the array at an angle corresponds to interchannel phase difference (IPD) data that lies on a wrapped line (i.e helix) in a circular-linear domain. Thus, the problem reduces to that of fitting helices to data that lies on a cylinder. However, outliers abound because of reverberation, noise, and signal overlap in the time-frequency domain, so we perform the clustering with a sequential variant of Random Sample Consensus (RANSAC). We show that this method can easily be applied to arrays with many microphones and that it is robust in reverberant experimental conditions.
Johannes Traa, Paris Smaragdis
ICASSP2
2013 Manifold Preserving Hierarchical Topic Models for Quantization and Approximation
abstract
We present two complementary topic models to address the analysis of mixture data lying on manifolds. First, we propose a quantization method with an additional mid-layer latent variable, which selects only data points that best preserve the manifold structure of the input data. In order to address the case of modeling all the in-between parts of that manifold using this reduced representation of the input, we introduce a new model that provides a manifold-aware interpolation method. We demonstrate the advantages of these models with experiments on the hand-written digit recognition and the speech source separation tasks.
Minje Kim 0001, Paris Smaragdis
ICML (3)2
2013 EMERALD: Characterization of emerging applications and algorithms for low-power devices
abstract
Compute-intensive applications are emerging in intelligent home, retail store and automotive industries. These applications are becoming more sophisticated with new features rich in audio, video, image, and machine learning capabilities that demand heavy computations. We present the EMERALD (EMERging Applications and algorithms for Low power Device) workload suite. We profile the workloads to show the hotspot functions that are candidates for hardware accelerators.
Chuanjun Zhang, Glenn G. Ko, Jungwook Choi, Shang-nien Tsai, Minje Kim 0001, Abner Guzmán-Rivera, Rob A. Rutenbar, Paris Smaragdis, Mi Sun Park, Narayanan Vijaykrishnan, Hongyi Xin, Onur Mutlu, Bin Li 0018, Li Zhao 0002
ISPASS8
2013 A Wrapped Kalman Filter for Azimuthal Speaker Tracking
abstract
We present the wrapped Kalman filter (WKF) for tracking the azimuth of a speaker with a compact, 3-channel microphone array. Traditional extended and unscented filters assume that the observation is a rotating vector in${\BBR^2}$. However, the azimuth inhabits a 1-D subspace: the unit circle. We model the state variable with a wrapped Gaussian distribution and show that this achieves a lower mean squared error than 2-D methods. We demonstrate the superior tracking performance of the WKF in simulated and real reverberant environments.
Johannes Traa, Paris Smaragdis
IEEE Signal Process. Lett.2
2013 Supervised and Unsupervised Speech Enhancement Using Nonnegative Matrix Factorization
abstract
Reducing the interference noise in a monaural noisy speech signal has been a challenging task for many years. Compared to traditional unsupervised speech enhancement methods, e.g., Wiener filtering, supervised approaches, such as algorithms based on hidden Markov models (HMM), lead to higher-quality enhanced speech signals. However, the main practical difficulty of these approaches is that for each noise type a model is required to be trained a priori. In this paper, we investigate a new class of supervised speech denoising algorithms using nonnegative matrix factorization (NMF). We propose a novel speech enhancement method that is based on a Bayesian formulation of NMF (BNMF). To circumvent the mismatch problem between the training and testing stages, we propose two solutions. First, we use an HMM in combination with BNMF (BNMF-HMM) to derive a minimum mean square error (MMSE) estimator for the speech signal with no information about the underlying noise type. Second, we suggest a scheme to learn the required noise BNMF model online, which is then used to develop an unsupervised speech enhancement system. Extensive experiments are carried out to investigate the performance of the proposed methods under different conditions. Moreover, we compare the performance of the developed algorithms with state-of-the-art speech enhancement schemes using various objective measures. Our simulations show that the proposed BNMF-based methods outperform the competing algorithms substantially.
Nasser Mohammadiha, Paris Smaragdis, Arne Leijon
IEEE Trans. Speech Audio Process.2
2012 Clustering and synchronizing multi-camera video via landmark cross-correlation
abstract
We propose a method to both identify and synchronize multi-camera video recordings within a large collection of video and/or audio files. Landmark-based audio fingerprinting is used to match multiple recordings of the same event together and time-synchronize each file within the groups. Compared to prior work, we offer improvements towards event identification and a new synchronization refinement method that resolves inconsistent estimates and allows non-overlapping content to be synchronized within larger groups of recordings. Furthermore, the audio fingerprinting-based synchronization is shown to be equivalent to an efficient and scalable time-difference-of-arrival method using cross-correlation performed on a non-linearly transformed signal.
Nicholas J. Bryan, Paris Smaragdis, Gautham J. Mysore
ICASSP2
2012 Singing-voice separation from monaural recordings using robust principal component analysis
abstract
Separating singing voices from music accompaniment is an important task in many applications, such as music information retrieval, lyric recognition and alignment. Music accompaniment can be assumed to be in a low-rank subspace, because of its repetition structure; on the other hand, singing voices can be regarded as relatively sparse within songs. In this paper, based on this assumption, we propose using robust principal component analysis for singing-voice separation from music accompaniment. Moreover, we examine the separation result by using a binary time-frequency masking method. Evaluations on the MIR-1K dataset show that this method can achieve around 1~1.4 dB higher GNSDR compared with two state-of-the-art approaches without using prior training or requiring particular features.
Po-Sen Huang, Scott Deeann Chen, Paris Smaragdis, Mark Hasegawa-Johnson
ICASSP3
2012 Noise-robust dynamic time warping using PLCA features
abstract
Conventional speech features, such as mel-frequency cepstral coefficients, tend to perform well in template matching systems, such as dynamic time warping, in low noise conditions. However, they tend to degrade in noisy environments. We propose a method of calculating features using the probabilistic latent component analysis (PLCA) framework. This framework models the speech and noise separately, leading to higher performance in noisy conditions than conventional methods. In this work, we compare our PLCA-based features with conventional features on the task of aligning a high-fidelity speech recording to a noisy speech recording, a scenario common in automatic dialogue replacement.
Brian King, Paris Smaragdis, Gautham J. Mysore
ICASSP2
2012 Following musical sources by example
abstract
In this paper we present a system that is capable of tracking the pitch and volume of a musical source by making use of training data. We show how we can use pitch-tagged training example sounds to construct a model of a target source, and then use that model to track such a source in unseen mixtures. We do so using a regularized decomposition approach that is designed to strive for semantic continuity in its estimates.
Paris Smaragdis, Gautham J. Mysore
ICASSP1
2012 Speech Enhancement by Online Non-negative Spectrogram Decomposition in Non-stationary Noise Environments
abstract
Classical single-channel speech enhancement algorithms have two convenient properties: they require pre-learning the noise model but not the speech model, and they work online. However, they often have difficulties in dealing with non-stationary noise sources. Source separation algorithms based on nonnegative spectrogram decompositions are capable of dealing with non-stationary noise, but do not possess the aforementioned properties. In this paper we present a novel algorithm that combines the advantages of both classical algorithms and non-negative spectrogram decomposition algorithms. Experiments show that it significantly outperforms four categories of classical algorithms in non-stationary noise environments.
Zhiyao Duan, Gautham J. Mysore, Paris Smaragdis
INTERSPEECH3
2012 The Markov selection model for concurrent speech recognition
Paris Smaragdis, Bhiksha Raj
Neurocomputing1
2011 An adaptive time-frequency resolution approach for Non-negative Matrix Factorization based single channel sound source separation
abstract
In this paper, we propose an adaptive time-frequency resolution approach for the single channel source separation problem. The aim is to improve the quality and intelligibility of the separated sources by adapting the time-frequency resolution of the analysis window to the characteristic of the signal under consideration. The results evaluated on a large test set show the improvements obtained by the proposed algorithm.
Serap Kirbiz, Paris Smaragdis
ICASSP2
2011 A non-negative approach to semi-supervised separation of speech from noise with the use of temporal dynamics
abstract
We present a semi-supervised source separation methodology to denoise speech by modeling speech as one source and noise as the other source. We model speech using the recently pro posed non-negative hidden Markov model, which uses multiple non-negative dictionaries and a Markov chain to jointly model spectral structure and temporal dynamics of speech. We perform separation of the speech and noise using the recently proposed non-negative factorial hidden Markov model. Although the speech model is learned from training data, the noise model is learned during the separation process and re quires no training data. We show that the proposed method achieves superior results to using non-negative spectrogram factorization, which ignores the non-stationarity and temporal dynamics of speech.
Gautham J. Mysore, Paris Smaragdis
ICASSP2
2011 Approximate nearest-subspace representations for sound mixtures
abstract
In this paper we present a novel approach to describe sound mixtures which is based on a geometric viewpoint. In this approach we extend the idea of a nearest-neighbor representation to address the case of superimposed sources. We show that in order to account for mixing effects we need to perform a search for nearest-subspaces, as opposed to nearest-neighbors. In order to reduce the excessive computational complexity of this search we present an efficient algorithm to solve this problem which amounts to a sparse coding approach. We demonstrate the efficacy of this algorithm by using it to separate mixtures of speech.
Paris Smaragdis
ICASSP1
2011 Preface
Martin Heckmann, Bhiksha Raj, Paris Smaragdis
Speech Commun.3
2010 Latent-variable decomposition based dereverberation of monaural and multi-channel signals
abstract
We present an algorithm to dereverberate single- and multi-channel audio recordings. The proposed algorithm models the magnitude spectrograms of clean audio signals as histograms drawn from a multinomial process. Spectrograms of reverberated signals are obtained as histograms of draws from the PDF of the sum of two random variables, one representing the spectrogram of clean speech and the second the frequency decomposition of the room response. The spectrogram of the clean signal is computed as a maximum-likelihood estimate from the spectrogram of reverberant speech using an EM algorithm. Experimental evaluations show that the proposed algorithm is able to greatly reduce the reverberation effects in even highly reverberant signals captured in auditoria and other open spaces.
Rita Singh, Bhiksha Raj, Paris Smaragdis
ICASSP3
2010 Editorial for the Special Issue on Signal Models and Representations of Musical and Environmental Sounds
abstract
The 23 papers in this special issue focus on signal models and representations of musical and environmental sounds.
Bertrand David 0001, Masataka Goto, Laurent Daudet, Paris Smaragdis
IEEE Trans. Speech Audio Process.4
2009 Relative pitch estimation of multiple instruments
abstract
We present an algorithm based on probabilistic latent component analysis and employ it for relative pitch estimation of multiple instruments in polyphonic music. A multilayered positive deconvolution is performed concurrently on mixture constant-Q transforms to obtain a relative pitch track and timbral signature for each instrument. Initial experimental results on mixtures of two instruments are quite promising and show high levels of accuracy.
Gautham J. Mysore, Paris Smaragdis
ICASSP2
2009 A Sparse Non-Parametric Approach for Single Channel Separation of Known Sounds
abstract
In this paper we present an algorithm for separating mixed sounds from a monophonic recording. Our approach makes use of training data which allows us to learn representations of the types of sounds that compose the mixture. In contrast to popular methods that attempt to extract com- pact generalizable models for each sound from training data, we employ the training data itself as a representation of the sources in the mixture. We show that mixtures of known sounds can be described as sparse com- binations of the training data itself, and in doing so produce significantly better separation results as compared to similar systems based on compact statistical models.
Paris Smaragdis, Madhusudana V. S. Shashanka, Bhiksha Raj
NIPS1
2009 User guided audio selection from complex sound mixtures
abstract
In this paper we present a novel interface for selecting sounds in audio mixtures. Traditional interfaces in audio editors provide a graphical representation of sounds which is either a waveform, or some variation of a time/frequency transform. Although with these representations a user might be able to visually identify elements of sounds in a mixture, they do not facilitate object-specific editing (e.g. selecting only the voice of a singer in a song). This interface uses audio guidance from a user in order to select a target sound within a mixture. The user is asked to vocalize (or otherwise sonically represent) the desired target sound, and an automatic process identifies and isolates the elements of the mixture that best relate to the user's input. This way of pointing to specific parts of an audio stream allows a user to perform audio selections which would have been infeasible otherwise.
Paris Smaragdis
UIST1
2009 Dynamic Range Extension Using Interleaved Gains
abstract
We present a methodology to sample signals in such a way so as to avoid the effects of signal clipping due to a limited dynamic range. We do so by attenuating a selective subset of the data before it gets sampled, so that if clipping is detected after the sampling process we can easily estimate the missing samples using the nonclipped samples that were attenuated. We show that under sparsity assumptions it is possible to reconstruct the clipped samples and recover a satisfactory representation of the original signal. We provide an analysis of the side effects of this process and show that on average when sampling signals with highly varying or unknown gain, we can guarantee a significantly lower potential for signal distortion and noise.
Paris Smaragdis
IEEE Trans. Speech Audio Process.1
2008 Sparse and shift-invariant feature extraction from non-negative data
abstract
In this paper we describe a technique that allows the extraction of multiple local shift-invariant features from analysis of non-negative data of arbitrary dimensionality. Our approach employs a probabilistic latent variable model with sparsity constraints. We demonstrate its utility by performing feature extraction in a variety of domains ranging from audio to images and video.
Paris Smaragdis, Bhiksha Raj, Madhusudana V. S. Shashanka
ICASSP1
2008 Speech denoising using nonnegative matrix factorization with priors
abstract
We present a technique for denoising speech using nonnegative matrix factorization (NMF) in combination with statistical speech and noise models. We compare our new technique to standard NMF and to a state-of-the-art Wiener filter implementation and show improvements in speech quality across a range of interfering noise types.
Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis, Ajay Divakaran
ICASSP3
2008 Regularized non-negative matrix factorization with temporal dependencies for speech denoising
abstract
We present a tecchnique for denoising speech using temporally regularized nonnegative matrix factorization (NMF). In previ-ous work [1], we used a regularized NMF update to impose structure within each audio frame. In this paper, we add frame-to-frame regularization across time and show that this additional regularization can also improve our speech denoising results. We evaluate our algorithm on a range of nonstationary noise types and outperform a state-of-the-art Wiener filter implemen-tation. Index Terms: speech enhancement, source separation, speech modeling, speech processing
Kevin W. Wilson, Bhiksha Raj, Paris Smaragdis
INTERSPEECH3
2007 Sensor and Data Systems, Audio-Assisted Cameras and Acoustic Doppler Sensors
abstract
In this chapter we present two technologies for sensing and surveillance -audio-assisted cameras and acoustic Doppler sensors for gait recognition.
Kaustubh Kalgaonkar, Paris Smaragdis, Bhiksha Raj
CVPR2
2007 Bandwidth Expansionwith a pólya URN Model
abstract
We present a new statistical technique for the estimation of the high frequency components (4-8 kHz) of speech signals from narrow-band (0-4 kHz) signals. The magnitude spectra of broadband speech are modelled as the outcome of a Polya Urn process, that represents the spectra as the histogram of the outcome of several draws from a mixture multinomial distribution over frequency indices. The multinomial distributions that compose this process are learnt from a corpus of broadband (0-8 kHz) speech. To estimate high-frequency components of narrow-band speech, its spectra are also modelled as the outcome of draws from a mixture-multinomial process that is composed of the learnt multinomials, where the counts of the indices of higher frequencies have been obscured. The obscured high-frequency components are then estimated as the expected number of draws of their indices from the mixture-multinomial. Experiments conducted on bandlimited signals derived from the WSJ corpus show that the proposed procedure is able to accurately estimate the high frequency components of these signals.
Bhiksha Raj, Rita Singh, Madhusudana V. S. Shashanka, Paris Smaragdis
ICASSP (4)4
2007 Sparse Overcomplete Decomposition for Single Channel Speaker Separation
abstract
We present an algorithm for separating multiple speakers from a mixed single channel recording. The algorithm is based on a model proposed by Raj and Smaragdis (2005). The idea is to extract certain characteristic spectra-temporal basis functions from training data for individual speakers and decompose the mixed signals as linear combinations of these learned bases. In other words, their model extracts a compact code of basis functions that can explain the space spanned by spectral vectors of a speaker. In our model, we generate a sparse-distributed code where we have more basis functions than the dimensionality of the space. We propose a probabilistic framework to achieve sparsity. Experiments show that the resulting sparse code better captures the structure in data and hence leads to better separation.
Madhusudana V. S. Shashanka, Bhiksha Raj, Paris Smaragdis
ICASSP (2)3
2007 A Framework for Secure Speech Recognition
abstract
We present an algorithm that enables privacy-preserving speech recognition transactions between multiple parties. We assume two commonplace scenarios. One being the case where one of two parties has private speech data to be transcribed and the other party has private models for speech recognition. And the other being that of one party having a speech model to be trained using private data of multiple other parties. In both of the above cases data privacy is desired from both the data and the model owners. In this paper we will show how such collaborations can be performed while ensuring no private data leaks using secure multiparty computations. In neither case will any party obtain information on other parties data. The protocols described herein can be used to construct rudimentary speech recognition systems and can be easily extended for arbitrary audio and speech processing.
Paris Smaragdis, Madhusudana V. S. Shashanka
ICASSP (4)1
2007 Sparse Overcomplete Latent Variable Decomposition of Counts Data
abstract
An important problem in many fields is the analysis of counts data to extract meaningful latent components. Methods like Probabilistic Latent Semantic Analysis (PLSA) and Latent Dirichlet Allocation (LDA) have been proposed for this purpose. However, they are limited in the number of components they can extract and also do not have a provision to control the expressiveness" of the extracted components. In this paper, we present a learning formulation to address these limitations by employing the notion of sparsity. We start with the PLSA framework and use an entropic prior in a maximum a posteriori formulation to enforce sparsity. We show that this allows the extraction of overcomplete sets of latent components which better characterize the data. We present experimental evidence of the utility of such representations."
Madhusudana V. S. Shashanka, Bhiksha Raj, Paris Smaragdis
NIPS3
2007 Convolutive Speech Bases and Their Application to Supervised Speech Separation
abstract
In this paper, we present a convolutive basis decomposition method and its application on simultaneous speakers separation from monophonic recordings. The model we propose is a convolutive version of the nonnegative matrix factorization algorithm. Due to the nonnegativity constraint this type of coding is very well suited for intuitively and efficiently representing magnitude spectra. We present results that reveal the nature of these basis functions and we introduce their utility in separating monophonic mixtures of known speakers
Paris Smaragdis
IEEE Trans. Speech Audio Process.1
2007 Position and Trajectory Learning for Microphone Arrays
abstract
In this paper, we tackle the problem of source localization by example. We present a methodology that allows a user to train a microphone array system using signals from a set of positions and trajectories and subsequently recall the localization information when presented with new input signals. To do so we present a new statistical model which is capable of accurately describing features from the cross spectra of the microphone signals so as to model the room responses from all positions of interest. We further extend this model to allow modeling of sequences of positions, thereby also enabling the learning and recognition of trajectories. Because of its learning nature this method provides practical advantages in setting up a microphone array, by not requiring favorable room acoustics, careful element positioning or uniformity of sensors. It also introduces an approach to localization which can be extended to other problems requiring models of transfer functions. We present tests on synthetic and real-world data and present the resulting recognition rates for a variety of situations
Paris Smaragdis, Petros Boufounos
IEEE Trans. Speech Audio Process.1
2007 A Framework for Secure Speech Recognition
abstract
In this paper, we present a process which enables privacy-preserving speech recognition transactions between two parties. We assume one party with private speech data and one party with private speech recognition models. Our goal is to enable these parties to perform a speech recognition task using their data, but without exposing their private information to each other. We will demonstrate how using secure multiparty computation principles we can construct a system where this transaction is possible, and how this system is computationally and securely correct. The protocols described herein can be used to construct a rudimentary speech recognition system and can easily be extended for arbitrary audio and speech processing.
Paris Smaragdis, Madhusudana V. S. Shashanka
IEEE Trans. Speech Audio Process.1
2006 Latent Dirichlet Decomposition for Single Channel Speaker Separation
abstract
We present an algorithm for the separation of multiple speakers from mixed single-channel recordings by latent variable decomposition of the speech spectrogram. We model each magnitude spectral vector in the short-time Fourier transform of a speech signal as the outcome of a discrete random process that generates frequency bin indices. The distribution of the process is modeled as a mixture of multinomial distributions, such that the mixture weights of the component multinomials vary from analysis window to analysis window. The component multinomials are assumed to be speaker specific and are learned from training signals for each speaker. We model the prior distribution of the mixture weights for each speaker as a Dirichlet distribution. The distributions representing magnitude spectral vectors for the mixed signal are decomposed into mixtures of the multinomials for all component speakers. The frequency distribution, i.e the spectrum for each speaker, is reconstructed from this decomposition
Bhiksha Raj, Madhusudana V. S. Shashanka, Paris Smaragdis
ICASSP (5)3
2006 Secure Sound Classification: Gaussian Mixture Models
abstract
We propose secure protocols for Gaussian mixture-based sound recognition. The protocols we describe allow varying levels of security between two collaborating parties. The case we examine consists of one party (Alice) providing data and other party (Bob) providing a recognition algorithm. We show that it is possible to have Bob apply his algorithm on Alice's data in such a way that the data and the recognition results will not be revealed to Bob thereby guaranteeing Alice's data privacy. Likewise we show that it is possible to organize the collaboration so that a reverse engineering of Bob's recognition algorithm cannot be performed by Alice. We show how Gaussian mixtures can be implemented in a secure manner using secure computation primitives implementing simple numerical operations and we demonstrate the process by showing how it can yield identical results to a non-secure computation while maintaining privacy
Madhusudana V. S. Shashanka, Paris Smaragdis
ICASSP (3)2
2005 Bandwidth expansion of narrowband speech using non-negative matrix factorization
abstract
In this paper, we present a novel technique for the estimation of the high frequency components (4-8kHz) of speech signals from narrow-band (0-4 kHz) signals using convolutive Non-Negative Matrix Factorisation (NMF). The proposed technique utilizes a brief recording of simultaneous broad band and narrow band signals from a target speaker to learn a set of broad-band non-negative bases for the speaker. The low-frequency components of these bases are used to determine how the high-frequency components must be combined in order to reconstruct the high-frequency components of new narrow-band signals from the speaker. Experiments reveal that the technique is able to reconstruct broadband sppech that is perceptually virtually indistinguishable from true broadband recordings.
Dhananjay Bansal, Bhiksha Raj, Paris Smaragdis
INTERSPEECH3
2005 Recognizing speech from simultaneous speakers
abstract
In this paper we present and evaluate factored methods for recognition of simultaneous speech from multiple speakers in single-channel recordings. Factored methods decompose the problem of jointly recognizing the speech from each of the speakers by separately recognizing the speech from each speaker. In order to achieve this, the signal components of the target speaker in each case must be enhanced in some manner. We do this in two ways: using an NMF-based speaker separation algorithm that generates separated spectra for each speaker, and a mask estimation method that generates spectral masks for each speaker that must be used in conjunction with a missing-feature method that can recognize speech from partial spectral data. Experiments on synthetic mixtures of signals from the Wall Street Journal corpus show that both approaches can greatly improve the recognition of the individual signals in the mixture.
Bhiksha Raj, Rita Singh, Paris Smaragdis
INTERSPEECH3
1998 Blind separation of convolved mixtures in the frequency domain
Paris Smaragdis
Neurocomputing1