EDBT 2026 Demo / reviewers in the wild / expert
Tuomas Virtanen
dblp:87/6879
· DBLP profile ↗
108ranked-venue papers
9as first author
23since 2021 · last 2025
0000-0002-4604-9729ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 67 · 7 first-author · 15 since 2021Artificial intelligence and machine learning · 54 · 4 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Gen-A: Generalizing Ambisonics Neural Encoding to Unseen Microphone ArraysabstractUsing deep neural networks (DNNs) for encoding of microphone array (MA) signals to the Ambisonics spatial audio format can surpass certain limitations of established conventional methods, but existing DNN-based methods need to be trained separately for each MA. This paper proposes a DNN-based method for Ambisonics encoding that can generalize to arbitrary MA geometries unseen during training. The method takes as inputs the MA geometry and MA signals and uses a multi-level encoder consisting of separate paths for geometry and signal data, where geometry features inform the signal encoder at each level. The method is validated in simulated anechoic and reverberant conditions with one and two sources. The results indicate improvement over conventional encoding across the whole frequency range for dry scenes, while for reverberant scenes the improvement is frequency-dependent. Mikko Heikkinen, Archontis Politis, Konstantinos Drossos, Tuomas Virtanen |
ICASSP | 4 |
| 2025 | A decade of DCASE: Achievements, practices, evaluations and future challengesabstractThis paper introduces briefly the history and growth of the Detection and Classification of Acoustic Scenes and Events (DCASE) challenge, workshop, research area and research community. Created in 2013 as a data evaluation challenge, DCASE has become a major research topic in the Audio and Acoustic Signal Processing area. Its success comes from a combination of factors: the challenge offers a large variety of tasks that are renewed each year; and the workshop offers a channel for dissemination of related work, engaging a young and dynamic community. At the same time, DCASE faces its own challenges, growing and expanding to different areas. One of the core principles of DCASE is open science and reproducibility: publicly available datasets, baseline systems, technical reports and workshop publications. While the DCASE challenge and workshop are independent of IEEE SPS, the challenge receives annual endorsement from the AASP TC, and the DCASE community contributes significantly to the ICASSP flagship conference and the success of SPS in many of its activities. Annamaria Mesaros, Romain Serizel, Toni Heittola, Tuomas Virtanen, Mark D. Plumbley |
ICASSP | 4 |
| 2025 | Inter-Speaker Relative Cues for Text-Guided Target Speech Extraction
Wang Dai, Archontis Politis, Tuomas Virtanen |
INTERSPEECH | 3 |
| 2025 | Attractor-Based Speech Separation of Multiple Utterances by Unknown Number of Speakers
Archontis Politis, Konstantinos Drossos, Tuomas Virtanen |
INTERSPEECH | 4 |
| 2025 | Computer Audition: From Task-Specific Machine Learning to Foundation ModelsabstractFoundation models (FMs) are increasingly spearheading recent advances on a variety of tasks that fall under the purview of computer audition—i.e., the use of machines to understand sounds. They feature several advantages over traditional pipelines: among others, the ability to consolidate multiple tasks in a single model, the option to leverage knowledge from other modalities, and the readily available interaction with human users. Naturally, these promises have created substantial excitement in the audio community and have led to a wave of early attempts to build new, generalpurpose FMs for audio. In the present contribution, we give an overview of computational audio analysis as it transitions from traditional pipelines toward auditory FMs. Our work highlights the key operating principles that underpin those models and showcases how they can accommodate multiple tasks that the audio community previously tackled separately. Andreas Triantafyllopoulos, Iosif Tsangko, Alexander Gebhard 0001, Annamaria Mesaros, Tuomas Virtanen, Björn W. Schuller |
Proc. IEEE | 5 |
| 2025 | Text-Based Audio Retrieval by Learning From Similarities Between Audio CaptionsabstractThis letter proposes to use similarities of audio captions for estimating audio-caption relevances to be used for training text-based audio retrieval systems. Current audio-caption datasets (e.g., Clotho) contain audio samples paired with annotated captions, but lack relevance information about audio samples and captions beyond the annotated ones. Besides, mainstream approaches (e.g., CLAP) usually treat the annotated pairs as positives and consider all other audio-caption combinations as negatives, assuming a binary relevance between audio samples and captions. To infer the relevance between audio samples and arbitrary captions, we propose a method that computes non-binary audio-caption relevance scores based on the textual similarities of audio captions. We measure textual similarities of audio captions by calculating the cosine similarity of their Sentence-BERT embeddings and then transform these similarities into audio-caption relevance scores using a logistic function, thereby linking audio samples through their annotated captions to all other captions in the dataset. To integrate the computed relevances into training, we employ a listwise ranking objective, where relevance scores are converted into probabilities of ranking audio samples for a given textual query. We show the effectiveness of the proposed method by demonstrating improvements in text-based audio retrieval compared to methods that use binary audio-caption relevances for training. Huang Xie, Khazar Khorrami, Okko Johannes Räsänen, Tuomas Virtanen |
IEEE Signal Process. Lett. | 4 |
| 2024 | Neural Ambisonics Encoding For Compact Irregular Microphone ArraysabstractAmbisonics encoding of microphone array signals can enable various spatial audio applications, such as virtual reality or telepresence, but it is typically designed for uniformly-spaced spherical microphone arrays. This paper proposes a method for Ambisonics encoding that uses a deep neural network (DNN) to estimate a signal transform from microphone inputs to Ambisonics signals. The approach uses a DNN consisting of a U-Net structure with a learnable preprocessing as well as a loss function consisting of mean average error, spatial correlation, and energy preservation components. The method is validated on two microphone arrays with regular and irregular shapes having four microphones, on simulated reverberant scenes with multiple sources. The results of the validation show that the proposed method can meet or exceed the performance of a conventional signal-independent Ambisonics encoder on a number of error metrics. Mikko Heikkinen, Archontis Politis, Tuomas Virtanen |
ICASSP | 3 |
| 2024 | Attention-Driven Multichannel Speech Enhancement in Moving Sound Source ScenariosabstractCurrent multichannel speech enhancement algorithms typically assume a stationary sound source, a common mismatch with reality that limits their performance in real-world scenarios. This paper focuses on attention-driven spatial filtering techniques designed for dynamic settings. Specifically, we study the application of linear and nonlinear attention-based methods for estimating time-varying spatial covariance matrices used to design the filters. We also investigate the direct estimation of spatial filters by attention-based methods without explicitly estimating spatial statistics. The clean speech clips from WSJ0 are employed for simulating speech signals of moving speakers in a reverberant environment. The experimental dataset is built by mixing the simulated speech signals with multichannel real noise from CHiME-3. Evaluation results show that the attention-driven approaches are robust and consistently outperform conventional spatial filtering approaches in both static and dynamic sound environments. Archontis Politis, Tuomas Virtanen |
ICASSP | 3 |
| 2024 | Dynamic Processing Neural Network Architecture for Hearing Loss CompensationabstractThis paper proposes neural networks for compensating sensorineural hearing loss. The aim of the hearing loss compensation task is to transform a speech signal to increase speech intelligibility after further processing by a person with a hearing impairment, which is modeled by a hearing loss model. We propose an interpretable model called dynamic processing network, which has a structure similar to band-wise dynamic compressor. The network is differentiable, and therefore allows to learn its parameters to maximize speech intelligibility. More generic models based on convolutional layers were tested as well. The performance of the tested architectures was assessed using spectro-temporal objective index (STOI) with hearing-threshold noise and hearing aid speech intelligibility (HASPI) metrics. The dynamic processing network gave a significant improvement of STOI and HASPI in comparison to popular compressive gain prescription rule Camfit. A large enough convolutional network could outperform the interpretable model with the cost of larger computational load. Finally, a combination of the dynamic processing network with convolutional neural network gave the best results in terms of STOI and HASPI. Szymon Drgas, Lars Bramslow, Archontis Politis, Gaurav Naithani, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Speaker Distance Estimation in Enclosures From Single-Channel AudioabstractDistance estimation from audio plays a crucial role in various applications, such as acoustic scene analysis, sound source localization, and room modeling. Most studies predominantly center on employing a classification approach, where distances are discretized into distinct categories, enabling smoother model training and achieving higher accuracy but imposing restrictions on the precision of the obtained sound source position. Towards this direction, in this paper we propose a novel approach for continuous distance estimation from audio signals using a convolutional recurrent neural network with an attention module. The attention mechanism enables the model to focus on relevant temporal and spectral features, enhancing its ability to capture fine-grained distance-related information. To evaluate the effectiveness of our proposed method, we conduct extensive experiments using audio recordings in controlled environments with three levels of realism (synthetic room impulse response, measured response with convolved speech, and real recordings) on four datasets (our synthetic dataset, QMULTIMIT, VoiceHome-2, and STARSS23). Experimental results show that the model achieves an absolute error of 0.11 meters in a noiseless synthetic scenario. Moreover, the results showed an absolute error of about 1.30 meters in the hybrid scenario. The algorithm's performance in the real scenario, where unpredictable environmental factors and noise are prevalent, yields an absolute error of approximately 0.50 meters. For reproducible research purposes we make model, code, and synthetic datasets available at https://github.com/michaelneri/audio-distance-estimation Michael Neri, Archontis Politis, Daniel Krause 0001, Marco Carli, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | On Negative Sampling for Contrastive Audio-Text RetrievalabstractThis paper investigates negative sampling for contrastive learning in the context of audio-text retrieval. The strategy for negative sampling refers to selecting negatives (either audio clips or textual descriptions) from a pool of candidates for a positive audio-text pair. We explore sampling strategies via model-estimated within-modality and cross-modality relevance scores for audio and text samples. With a constant training setting on the retrieval system from [1], we study eight sampling strategies, including hard and semi-hard negative sampling. Experimental results show that retrieval performance varies dramatically among different strategies. Particularly, by selecting semi-hard negatives with cross-modality scores, the retrieval system gains improved performance in both text-to-audio and audio-to-text retrieval. Besides, we show that feature collapse occurs while sampling hard negatives with cross-modality scores. Huang Xie, Okko Johannes Räsänen, Tuomas Virtanen |
ICASSP | 3 |
| 2023 | Few-shot Class-incremental Audio Classification Using Adaptively-refined PrototypesabstractNew classes of sounds constantly emerge with a few samples, making it challenging for models to adapt to dynamic acoustic environments. This challenge motivates us to address the new problem of few-shot class-incremental audio classification. This study aims to enable a model to continuously recognize new classes of sounds with a few training samples of new classes while remembering the learned ones. To this end, we propose a method to generate discriminative prototypes and use them to expand the model's classifier for recognizing sounds of new and learned classes. The model is first trained with a random episodic training strategy, and then its backbone is used to generate the prototypes. A dynamic relation projection module refines the prototypes to enhance their discriminability. Results on two datasets (derived from the corpora of Nsynth and FSD-MIX-CLIPS) show that the proposed method exceeds three state-of-the-art methods in average accuracy and performance dropping rate. Wei Xie 0013, Yanxiong Li, Qianhua He, Wenchang Cao, Tuomas Virtanen |
INTERSPEECH | 5 |
| 2023 | STARSS23: An Audio-Visual Dataset of Spatial Recordings of Real Scenes with Spatiotemporal Annotations of Sound EventsabstractWhile direction of arrival (DOA) of sound events is generally estimated from multichannel audio data recorded in a microphone array, sound events usually derive from visually perceptible source objects, e.g., sounds of footsteps come from the feet of a walker. This paper proposes an audio-visual sound event localization and detection (SELD) task, which uses multichannel audio and video information to estimate the temporal activation and DOA of target sound events. Audio-visual SELD systems can detect and localize sound events using signals from a microphone array and audio-visual correspondence. We also introduce an audio-visual dataset, Sony-TAu Realistic Spatial Soundscapes 2023 (STARSS23), which consists of multichannel audio data recorded with a microphone array, video data, and spatiotemporal annotation of sound events. Sound scenes in STARSS23 are recorded with instructions, which guide recording participants to ensure adequate activity and occurrences of sound events. STARSS23 also serves human-annotated temporal activation labels and human-confirmed DOA labels, which are based on tracking results of a motion capture system. Our benchmark results demonstrate the benefits of using visual object positions in audio-visual SELD tasks. The data is available at https://zenodo.org/record/7880637. Kazuki Shimada, Archontis Politis, Parthasaarathy Sudarsanam, Daniel Krause 0001, Kengo Uchida, Sharath Adavanne, Aapo Hakala, Yuichiro Koyama, Naoya Takahashi, Shusuke Takahashi, Tuomas Virtanen, Yuki Mitsufuji |
NeurIPS | 11 |
| 2022 | Unsupervised Audio-Caption Aligning Learns Correspondences Between Individual Sound Events and Textual PhrasesabstractWe investigate unsupervised learning of correspondences between sound events and textual phrases through aligning audio clips with textual captions describing the content of a whole audio clip. We align originally unaligned and unannotated audio clips and their captions by scoring the similarities between audio frames and words, as encoded by modality-specific encoders and using a ranking-loss criterion to optimize the model. After training, we obtain clip-caption similarity by averaging frame-word similarities and estimate event-phrase correspondences by calculating frame-phrase similarities. We evaluate the method with two cross-modal tasks: audio-caption retrieval, and phrase-based sound event detection (SED). Experimental results show that the proposed method can globally associate audio clips with captions as well as locally learn correspondences between individual sound events and textual phrases in an unsupervised manner. Huang Xie, Okko Johannes Räsänen, Konstantinos Drossos, Tuomas Virtanen |
ICASSP | 4 |
| 2022 | Domestic Activity Clustering from Audio via Depthwise Separable Convolutional Autoencoder NetworkabstractAutomatic estimation of domestic activities from audio can be used to solve many problems, such as reducing the labor cost for nursing the elderly people. This study focuses on solving the problem of domestic activity clustering from audio. The target of domestic activity clustering is to cluster audio clips which belong to the same category of domestic activity into one cluster in an unsupervised way. In this paper, we propose a method of domestic activity clustering using a depthwise separable convolutional autoencoder network. In the proposed method, initial embeddings are learned by the depthwise separable convolutional autoencoder, and a clustering-oriented loss is designed to jointly optimize embedding refinement and cluster assignment. Different methods are evaluated on a public dataset (a derivative of the SINS dataset) used in the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) in 2018. Our method obtains the normalized mutual information (NMI) score of 54.46%, and the clustering accuracy (CA) score of 63.64%, and outperforms state-of-the-art methods in terms of NMI and CA. In addition, both computational complexity and memory requirement of our method is lower than that of previous deep-model-based methods. Codes: https://github.com/vinceasvp/domestic-activity-clustering-from-audio Yanxiong Li, Wenchang Cao, Konstantinos Drossos, Tuomas Virtanen |
MMSP | 4 |
| 2022 | Subjective Evaluation of Deep Neural Network Based Speech Enhancement Systems in Real-World ConditionsabstractSubjective evaluation results for two low-latency deep neural networks (DNN) are compared to a matured version of a traditional Wiener-filter based noise suppressor. The target use-case is real-world single-channel speech enhancement applications, e.g., communications. Real-world recordings consisting of additive stationary and non-stationary noise types are included. The evaluation is divided into four outcomes: speech quality, noise transparency, speech intelligibility or listening effort, and noise level w.r.t. speech. It is shown that DNNs improve noise suppression in all conditions in comparison to the traditional Wiener-filter baseline without major degradation in speech quality and noise transparency while maintaining speech intelligibility better than the baseline. Gaurav Naithani, Kirsi Pietilä, Riitta Niemistö, Erkki Paajanen, Tero Takala, Tuomas Virtanen |
MMSP | 6 |
| 2021 | Learning Contextual Tag Embeddings for Cross-Modal Alignment of Audio and TagsabstractSelf-supervised audio representation learning offers an attractive alternative for obtaining generic audio embeddings, capable to be employed into various downstream tasks. Published approaches that consider both audio and words/tags associated with audio do not employ text processing models that are capable to generalize to tags unknown during training. In this work we propose a method for learning audio representations using an audio autoencoder (AAE), a general word embed-dings model (WEM), and a multi-head self-attention (MHA) mechanism. MHA attends on the output of the WEM, pro-viding a contextualized representation of the tags associated with the audio, and we align the output of MHA with the out-put of the encoder of AAE using a contrastive loss. We jointly optimize AAE and MHA and we evaluate the audio representations (i.e. the output of the encoder of AAE) by utilizing them in three different downstream tasks, namely sound, music genre, and music instrument classification. Our results show that employing multi-head self-attention with multiple heads in the tag-based network can induce better learned audio representations. Xavier Favory, Konstantinos Drossos, Tuomas Virtanen, Xavier Serra |
ICASSP | 3 |
| 2021 | A Curated Dataset of Urban Scenes for Audio-Visual Scene AnalysisabstractThis paper introduces a curated dataset of urban scenes for audio-visual scene analysis which consists of carefully selected and recorded material. The data was recorded in multiple European cities, using the same equipment, in multiple locations for each scene, and is openly available. We also present a case study for audio-visual scene recognition and show that joint modeling of audio and visual modalities brings significant performance gain compared to state of the art uni-modal systems. Our approach obtained an 84.8% ac-curacy compared to 75.8% for the audio-only and 68.4% for the video-only equivalent systems. Shanshan Wang 0010, Annamaria Mesaros, Toni Heittola, Tuomas Virtanen |
ICASSP | 4 |
| 2021 | Zero-Shot Audio Classification with Factored Linear and Nonlinear Acoustic-Semantic ProjectionsabstractIn this paper, we study zero-shot learning in audio classification through factored linear and nonlinear acoustic-semantic projections between audio instances and sound classes. Zero-shot learning in audio classification refers to classification problems that aim at recognizing audio instances of sound classes, which have no available training data but only semantic side information. In this paper, we address zero-shot learning by employing factored linear and nonlinear acoustic-semantic projections. We develop factored linear projections by applying rank decomposition to a bilinear model, and use nonlinear activation functions, such as tanh, to model the non-linearity between acoustic embeddings and semantic embeddings. Compared with the prior bilinear model, experimental results show that the proposed projection methods are effective for improving classification performance of zero-shot learning in audio classification. Huang Xie, Okko Johannes Räsänen, Tuomas Virtanen |
ICASSP | 3 |
| 2021 | Towards Sonification in Multimodal and User-friendlyExplainable Artificial IntelligenceabstractWe are largely used to hearing explanations. For example, if someone thinks you are sad today, they might reply to your “why?” with “because you were so Hmmmmm-mmm-mmm”. Today’s Artificial Intelligence (AI), however, is – if at all – largely providing explanations of decisions in a visual or textual manner. While such approaches are good for communication via visual media such as in research papers or screens of intelligent devices, they may not always be the best way to explain; especially when the end user is not an expert. In particular, when the AI’s task is about Audio Intelligence, visual explanations appear less intuitive than audible, sonified ones. Sonification has also great potential for explainable AI (XAI) in systems that deal with non-audio data – for example, because it does not require visual contact or active attention of a user. Hence, sonified explanations of AI decisions face a challenging, yet highly promising and pioneering task. That involves incorporating innovative XAI algorithms to allow pointing back at the learning data responsible for decisions made by an AI, and to include decomposition of the data to identify salient aspects. It further aims to identify the components of the preprocessing, feature representation, and learnt attention patterns that are responsible for the decisions. Finally, it targets decision-making at the model-level, to provide a holistic explanation of the chain of processing in typical pattern recognition problems from end-to-end. Sonified AI explanations will need to unite methods for sonification of the identified aspects that benefit decisions, decomposition and recomposition of audio to sonify which parts in the audio were responsible for the decision, and rendering attention patterns and salient feature representations audible. Benchmarking sonified XAI is challenging, as it will require a comparison against a backdrop of existing, state-of-the-art visual and textual alternatives, as well as synergistic complementation of all modalities in user evaluations. Sonified AI explanations will need to target different user groups to allow personalisation of the sonification experience for different user needs, to lead to a major breakthrough in comprehensibility of AI via hearing how decisions are made, hence supporting tomorrow’s humane AI’s trustability. Here, we introduce and motivate the general idea, and provide accompanying considerations including milestones of realisation of sonifed XAI and foreseeable risks. Björn W. Schuller, Tuomas Virtanen, Maria Riveiro 0001, Georgios Rizos, Jing Han 0010, Annamaria Mesaros, Konstantinos Drossos |
ICMI | 2 |
| 2021 | Joint speaker separation and recognition using non-negative matrix deconvolution with adaptive dictionary
Szymon Drgas, Tuomas Virtanen |
Comput. Speech Lang. | 2 |
| 2021 | Overview and Evaluation of Sound Event Localization and Detection in DCASE 2019abstractSound event localization and detection is a novel area of research that emerged from the combined interest of analyzing the acoustic scene in terms of the spatial and temporal activity of sounds of interest. This paper presents an overview of the first international evaluation on sound event localization and detection, organized as a task of the DCASE 2019 Challenge. A large-scale realistic dataset of spatialized sound events was generated for the challenge, to be used for training of learning-based approaches, and for evaluation of the submissions in an unlabeled subset. The overview presents in detail how the systems were evaluated and ranked and the characteristics of the best-performing systems. Common strategies in terms of input features, model architectures, training approaches, exploitation of prior knowledge, and data augmentation are discussed. Since ranking in the challenge was based on individually evaluating localization and event classification performance, part of the overview focuses on presenting metrics for the joint measurement of the two, together with a reevaluation of submissions using these new metrics. The new analysis reveals submissions that performed better on the joint task of detecting the correct type of event close to its original location than some of the submissions that were ranked higher in the challenge. Consequently, ranking of submissions which performed strongly when evaluated separately on detection or localization, but not jointly on both, was affected negatively. Archontis Politis, Annamaria Mesaros, Sharath Adavanne, Toni Heittola, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | Zero-Shot Audio Classification Via Semantic EmbeddingsabstractIn this paper, we study zero-shot learning in audio classification via semantic embeddings extracted from textual labels and sentence descriptions of sound classes. Our goal is to obtain a classifier that is capable of recognizing audio instances of sound classes that have no available training samples, but only semantic side information. We employ a bilinear compatibility framework to learn an acoustic-semantic projection between intermediate-level representations of audio instances and sound classes, i.e., acoustic embeddings and semantic embeddings. We use VGGish to extract deep acoustic embeddings from audio clips, and pre-trained language models (Word2Vec, GloVe, BERT) to generate either label embeddings from textual labels or sentence embeddings from sentence descriptions of sound classes. Audio classification is performed by a linear compatibility function that measures how compatible an acoustic embedding and a semantic embedding are. We evaluate the proposed method on a small balanced dataset ESC-50 and a large-scale unbalanced audio subset of AudioSet. The experimental results show that classification performance is significantly improved by involving sound classes that are semantically close to the test classes in training. Meanwhile, we demonstrate that both label embeddings and sentence embeddings are useful for zero-shot learning. Classification performance is improved by concatenating label/sentence embeddings generated with different language models. With their hybrid concatenations, the results are improved further. Huang Xie, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Clotho: an Audio Captioning DatasetabstractAudio captioning is the novel task of general audio content description using free text. It is an intermodal translation task (not speech-to-text), where a system accepts as an input an audio signal and outputs the textual description (i.e. the caption) of that signal. In this paper we present Clotho, a dataset for audio captioning consisting of 4981 audio samples of 15 to 30 seconds duration and 24 905 captions of eight to 20 words length, and a baseline method to provide initial results. Clotho is built with focus on audio content and caption diversity, and the splits of the data are not hampering the training or evaluation of methods. All sounds are from the Freesound platform, and captions are crowdsourced using Amazon Mechanical Turk and annotators from English speaking countries. Unique words, named entities, and speech transcription are removed with post-processing. Clotho is freely available online1. Konstantinos Drossos, Samuel Lipping, Tuomas Virtanen |
ICASSP | 3 |
| 2020 | Sound Event Detection Via Dilated Convolutional Recurrent Neural NetworksabstractConvolutional recurrent neural networks (CRNNs) have achieved state-of-the-art performance for sound event detection (SED). In this paper, we propose to use a dilated CRNN, namely a CRNN with a dilated convolutional kernel, as the classifier for the task of SED. We investigate the effectiveness of dilation operations which provide a CRNN with expanded receptive fields to capture long temporal context without increasing the amount of CRNN's parameters. Compared to the classifier of the baseline CRNN, the classifier of the dilated CRNN obtains a maximum increase of 1.9%, 6.3% and 2.5% at F1 score and a maximum decrease of 1.7%, 4.1% and 3.9% at error rate (ER), on the publicly available audio corpora of the TUTSED Synthetic 2016, the TUT Sound Event 2016 and the TUT Sound Event 2017, respectively. Yanxiong Li, Mingle Liu, Konstantinos Drossos, Tuomas Virtanen |
ICASSP | 4 |
| 2020 | Sound Event Detection with Depthwise Separable and Dilated ConvolutionsabstractState-of-the-art sound event detection (SED) methods usually employ a series of convolutional neural networks (CNNs) to extract useful features from the input audio signal, and then recurrent neural networks (RNNs) to model longer temporal context in the extracted features. The number of the channels of the CNNs and size of the weight matrices of the RNNs have a direct effect on the total amount of parameters of the SED method, which is to a couple of millions. Additionally, the usually long sequences that are used as an input to an SED method along with the employment of an RNN, introduce implications like increased training time, difficulty at gradient flow, and impeding the parallelization of the SED method. To tackle all these problems, we propose the replacement of the CNNs with depthwise separable convolutions and the replacement of the RNNs with dilated convolutions. We compare the proposed method to a baseline convolutional neural network on a SED task, and achieve a reduction of the amount of parameters by 85% and average training time per epoch by 78%, and an increase the average frame-wise F1score and reduction of the average error rate by 4.6% and 3.8%, respectively. Konstantinos Drossos, Stylianos I. Mimilakis, Shayan Gharib, Yanxiong Li, Tuomas Virtanen |
IJCNN | 5 |
| 2020 | Robust Audio-Based Vehicle Counting in Low-to-Moderate Traffic FlowabstractThe paper presents a method for audio-based vehicle counting (VC) in low-to-moderate traffic using one-channel sound. We formulate VC as a regression problem, i.e., we predict the distance between a vehicle and the microphone. Minima of the proposed distance function correspond to vehicles passing by the microphone. VC is carried out via local minima detection in the predicted distance. We propose to set the minima detection threshold at a point where the probabilities of false positives and false negatives coincide so they statistically cancel each other in total vehicle number. The method is trained and tested on a traffic-monitoring dataset comprising 422 short, 20-second one-channel sound files with a total of 1421 vehicles passing by the microphone. Relative VC error in a traffic location not used in the training is below 2% within a wide range of detection threshold values. Experimental results show that the regression accuracy in noisy environments is improved by introducing a novel high-frequency power feature. Slobodan Djukanovic, Jiri Matas, Tuomas Virtanen |
IV | 3 |
| 2020 | Depthwise Separable Convolutions Versus Recurrent Neural Networks for Monaural Singing Voice SeparationabstractRecent approaches for music source separation are almost exclusively based on deep neural networks, mostly employing recurrent neural networks (RNNs). Although RNNs are in many cases superior than other types of deep neural networks for sequence processing, they are known to have specific difficulties in training and parallelization, especially for the typically long sequences encountered in music source separation. In this paper we present a use-case of replacing RNNs with depth-wise separable (DWS) convolutions, which are a lightweight and faster variant of the typical convolutions. We focus on singing voice separation, employing an RNN architecture, and we replace the RNNs with DWS convolutions (DWS-CNNs). We conduct an ablation study and examine the effect of the number of channels and layers of DWS-CNNs on the source separation performance, by utilizing the standard metrics of signal-to-artifacts, signal-to-interference, and signal-to-distortion ratio. Our results show that by replacing RNNs with DWS-CNNs yields an improvement of 1.20, 0.06, 0.37 dB, respectively, while using only 20.57% of the amount of parameters of the RNN architecture. Pyry Pyykkönen, Stylianos I. Mimilakis, Konstantinos Drossos, Tuomas Virtanen |
MMSP | 4 |
| 2020 | Online Spectrogram Inversion for Low-Latency Audio Source SeparationabstractAudio source separation is usually achieved by estimating the short-time Fourier transform (STFT) magnitude of each source, and then applying a spectrogram inversion algorithm to retrieve time-domain signals. In particular, the multiple input spectrogram inversion (MISI) algorithm has been exploited successfully in several recent works. However, this algorithm suffers from two drawbacks, which we address in this letter. First, it has originally been introduced in a heuristic fashion: we propose here a rigorous optimization framework in which MISI is derived, thus proving the convergence of this algorithm. Besides, while MISI operates offline, we propose here an online version of MISI called oMISI, which is suitable for low-latency source separation, an important requirement for e.g., hearing aids applications. oMISI also allows one to use alternative phase initialization schemes exploiting the temporal structure of audio signals. Experiments conducted on a speech separation task show that oMISI performs as well as its offline counterpart, thus demonstrating its potential for real-time source separation. Paul Magron, Tuomas Virtanen |
IEEE Signal Process. Lett. | 2 |
| 2020 | Active Learning for Sound Event DetectionabstractThis article proposes an active learning system for sound event detection (SED). It aims at maximizing the accuracy of a learned SED model with limited annotation effort. The proposed system analyzes an initially unlabeled audio dataset, from which it selects sound segments for manual annotation. The candidate segments are generated based on a proposed change point detection approach, and the selection is based on the principle of mismatch-first farthest-traversal. During the training of SED models, recordings are used as training inputs, preserving the long-term context for annotated segments. The proposed system clearly outperforms reference methods in the two datasets used for evaluation (TUT Rare Sound 2017 and TAU Spatial Sound 2019). Training with recordings as context outperforms training with only annotated segments. Mismatch-first farthest-traversal outperforms reference sample selection methods based on random sampling and uncertainty sampling. Remarkably, the required annotation effort can be greatly reduced on the dataset where target sound events are rare: by annotating only 2% of the training data, the achieved SED performance is similar to annotating all the training data. Shuyang Zhao, Toni Heittola, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Sound Event Envelope Estimation in Polyphonic MixturesabstractSound event detection is the task of identifying automatically the presence and temporal boundaries of sound events within an input audio stream. In the last years, deep learning methods have established themselves as the state-of-the-art approach for the task, using binary indicators during training to denote whether an event is active or inactive. However, such binary activity indicators do not fully describe the events, and estimating the envelope of the sounds could provide more precise modeling of their activity. This paper proposes to estimate the amplitude envelopes of target sound event classes in polyphonic mixtures. For training, we use the amplitude envelopes of the target sounds, calculated from mixture signals and, for comparison, from their isolated counterparts. The model is then used to perform envelope estimation and sound event detection. Results show that the envelope estimation allows good modeling of the sounds activity, with detection results comparable to current state-of-the art. Irene Martín-Morató, Annamaria Mesaros, Toni Heittola, Tuomas Virtanen, Maximo Cobos, Francesc J. Ferri |
ICASSP | 4 |
| 2019 | Low-latency Deep Clustering for Speech SeparationabstractThis paper proposes a low algorithmic latency adaptation of the deep clustering approach to speaker-independent speech separation. It consists of three parts: a) the usage of long-short-term-memory (LSTM) networks instead of their bidirectional variant used in the original work, b) using a short synthesis window (here 8 ms) required for low-latency operation, and, c) using a buffer in the beginning of audio mixture to estimate cluster centres corresponding to constituent speakers which are then utilized to separate speakers within the rest of the signal. The buffer duration would serve as an initialization phase after which the system is capable of operating with 8 ms algorithmic latency. We evaluate our proposed approach on two-speaker mixtures from Wall Street Journal (WSJ0) corpus. We observe that the use of LSTM yields around one dB lower SDR as compared to the baseline bidirectional LSTM in terms of source to distortion ratio (SDR). Moreover, using an 8 ms synthesis window instead of 32 ms degrades the separation performance by around 2.1 dB as compared to the baseline. Finally, we also report separation performance with different buffer durations noting that separation can be achieved even for buffer duration as low as 300 ms. Shanshan Wang 0010, Gaurav Naithani, Tuomas Virtanen |
ICASSP | 3 |
| 2019 | Detection of Typical Pronunciation Errors in Non-native English Speech Using Convolutional Recurrent Neural NetworksabstractA machine learning method for the automatic detection of pronunciation errors made by non-native speakers of English is proposed. It consists of training word-specific binary classifiers on a collected dataset of isolated words with possible pronunciation errors, typical for Finnish native speakers. The classifiers predict whether the typical error is present in the given word utterance. They operate on sequences of acoustic features, extracted from consecutive frames of an audio recording of a word utterance. The proposed architecture includes a convolutional neural network, a recurrent neural network, or a combination of the two. The optimal topology and hyperparameters are obtained in a Bayesian optimisation setting using a tree-structured Parzen estimator. A dataset of 80 words uttered naturally by 120 speakers is collected. The performance of the proposed system, evaluated on a well-represented subset of the dataset, shows that it is capable of detecting pronunciation errors in most of the words (46/49) with high accuracy (mean accuracy gain over the zero rule 12.21 percent points). Aleksandr Diment, Eemi Fagerlund, Adrian Benfield, Tuomas Virtanen |
IJCNN | 4 |
| 2019 | Complex ISNMF: A Phase-Aware Model for Monaural Audio Source SeparationabstractThis paper introduces a phase-aware probabilistic model for audio source separation. Classical source models in the short-time Fourier transform domain use circularly-symmetric Gaussian or Poisson random variables. This is equivalent to assuming that the phase of each source is uniformly distributed, which is not suitable for exploiting the underlying structure of the phase. Drawing on preliminary works, we introduce here a Bayesian anisotropic Gaussian source model in which the phase is no longer uniform. Such a model permits us to favor a phase value that originates from a signal model through a Markov chain prior structure. The variance of the latent variables are structured with nonnegative matrix factorization (NMF). The resulting model is called complex Itakura-Saito NMF (ISNMF) since it generalizes the ISNMF model to the case of nonisotropic variables. It combines the advantages of ISNMF, which uses a distortion measure adapted to audio and yields a set of estimates which preserve the overall energy of the mixture, and of complex NMF, which enables one to account for some phase constraints. We derive a generalized expectation-maximization algorithm to estimate the model parameters. Experiments conducted on a musical source separation task in a semiinformed setting show that the proposed approach outperforms state-of-the-art phase-aware separation techniques. Paul Magron, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Sound Event Detection in the DCASE 2017 ChallengeabstractEach edition of the challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) contained several tasks involving sound event detection in different setups. DCASE 2017 presented participants with three such tasks, each having specific datasets and detection requirements: Task 2, in which target sound events were very rare in both training and testing data, Task 3 having overlapping events annotated in real-life audio, and Task 4, in which only weakly labeled data were available for training. In this paper, we present three tasks, including the datasets and baseline systems, and analyze the challenge entries for each task. We observe the popularity of methods using deep neural networks, and the still widely used mel frequency-based representations, with only few approaches standing out as radically different. Analysis of the systems behavior reveals that task-specific optimization has a big role in producing good performance; however, often this optimization closely follows the ranking metric, and its maximization/minimization does not result in universally good performance. We also introduce the calculation of confidence intervals based on a jackknife resampling procedure to perform statistical analysis of the challenge results. The analysis indicates that while the 95% confidence intervals for many systems overlap, there are significant differences in performance between the top systems and the baseline for all tasks. Annamaria Mesaros, Aleksandr Diment, Benjamin Elizalde, Toni Heittola, Emmanuel Vincent 0001, Bhiksha Raj, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 7 |
| 2019 | Analysis of an efficient parallel implementation of active-set Newton algorithm
Pablo San Juan Sebastián, Tuomas Virtanen, Víctor M. García 0001, Antonio M. Vidal |
J. Supercomput. | 2 |
| 2018 | Bayesian Anisotropic Gaussian Model for Audio Source SeparationabstractIn audio source separation applications, it is common to model the sources as circular-symmetric Gaussian random variables, which is equivalent to assuming that the phase of each source is uniformly distributed. In this paper, we introduce an anisotropic Gaussian source model in which both the magnitude and phase parameters are modeled as random variables. In such a model, it becomes possible to promote a phase value that originates from a signal model and to adjust the relative importance of this underlying model-based phase constraint. We conduct Bayesian inference of the model through the derivation of an expectation-maximization algorithm for estimating the parameters. Experiments conducted on realistic music songs for a monaural source separation task, in an scenario where the variance parameters are assumed known, show that the proposed approach outperforms state-of-the-art techniques. Paul Magron, Tuomas Virtanen |
ICASSP | 2 |
| 2018 | Monaural Singing Voice Separation with Skip-Filtering Connections and Recurrent Inference of Time-Frequency MaskabstractSinging voice separation based on deep learning relies on the usage of time-frequency masking. In many cases the masking process is not a learnable function or is not encapsulated into the deep learning optimization. Consequently, most of the existing methods rely on a post processing step using the generalized Wiener filtering. This work proposes a method that learns and optimizes (during training) a source-dependent mask and does not need the aforementioned post processing step. We introduce a recurrent inference algorithm, a sparse transformation step to improve the mask generation process, and a learned denoising filter. Obtained results show an increase of 0.49 dB for the signal to distortion ratio and 0.30 dB for the signal to interference ratio, compared to previous state-of-the-art approaches for monaural singing voice separation. Stylianos I. Mimilakis, Konstantinos Drossos, João Felipe Santos, Gerald Schuller, Tuomas Virtanen, Yoshua Bengio |
ICASSP | 5 |
| 2018 | Estimation of Time-Varying Room Impulse Responses of Multiple Sound Sources from Observed Mixture and Isolated Source SignalsabstractThis paper proposes a method for online estimation of time-varying room impulse responses (RIR) between multiple isolated sound sources and a far-field mixture. The algorithm is formulated as adaptive convolutive filtering in short-time Fourier transform (STFT) domain. We use the recursive least squares (RLS) algorithm for estimating the filter parameters due to its fast convergence rate, which is required for modeling rapidly changing RIRs of moving sound sources. The proposed method allows separation of reverberated sources from the far-field mixture given that their close-field signals are available. The evaluation is based on measuring unmixing performance (removal of reverberated source) using objective separation criteria calculated between the ground truth recording of the preserved sources and the unmixing result obtained with the proposed algorithm. We compare online and offline formulations for the RIR estimation and also provide evaluation with blind source separation algorithm only operating on the mixture signal. Joonas Nikunen, Tuomas Virtanen |
ICASSP | 2 |
| 2018 | Multichannel Sound Event Detection Using 3D Convolutional Neural Networks for Learning Inter-channel FeaturesabstractIn this paper, we propose a stacked convolutional and recurrent neural network (CRNN) with a 3D convolutional neural network (CNN) in the first layer for the multichannel sound event detection (SED) task. The 3D CNN enables the network to simultaneously learn the inter-and intra-channel features from the input multichannel audio. In order to evaluate the proposed method, multichannel audio datasets with different number of overlapping sound sources are synthesized. Each of this dataset has a four-channel first-order Ambisonic, binaural, and single-channel versions, on which the performance of SED using the proposed method are compared to study the potential of SED using multichannel audio. A similar study is also done with the binaural and single-channel versions of the real-life recording TUT-SED 2017 development dataset. The proposed method learns to recognize overlapping sound events from multichannel features faster and performs better SED with a fewer number of training epochs. The results show that on using multichannel Ambisonic audio in place of single-channel audio we improve the overall F-score by 7.5%, overall error rate by 10% and recognize 15.6% more sound events in time frames with four overlapping sound sources. Sharath Adavanne, Archontis Politis, Tuomas Virtanen |
IJCNN | 3 |
| 2018 | End-to-End Polyphonic Sound Event Detection Using Convolutional Recurrent Neural Networks with Learned Time-Frequency Representation InputabstractSound event detection systems typically consist of two stages: extracting hand-crafted features from the raw audio waveform, and learning a mapping between these features and the target sound events using a classifier. Recently, the focus of sound event detection research has been mostly shifted to the latter stage using standard features such as mel spectrogram as the input for classifiers such as deep neural networks. In this work, we utilize end-to-end approach and propose to combine these two stages in a single deep neural network classifier. The feature extraction over the raw waveform is conducted by a feedforward layer block, whose parameters are initialized to extract the time-frequency representations. The feature extraction parameters are updated during training, resulting with a representation that is optimized for the specific task. This feature extraction block is followed by (and jointly trained with) a convolutional recurrent network, which has recently given state-of-the-art results in many sound recognition tasks. The proposed system does not outperform a convolutional recurrent network with fixed hand-crafted features. The final magnitude spectrum characteristics of the feature extraction block parameters indicate that the most relevant information for the given task is contained in 0 - 3 kHz frequency range, and this is also supported by the empirical results on the SED performance. Emre Çakir, Tuomas Virtanen |
IJCNN | 2 |
| 2018 | MaD TwinNet: Masker-Denoiser Architecture with Twin Networks for Monaural Sound Source SeparationabstractMonaural singing voice separation task focuses on the prediction of the singing voice from a single channel music mixture signal. Current state of the art (SOTA) results in monaural singing voice separation are obtained with deep learning based methods. In this work we present a novel recurrent neural approach that learns long-term temporal patterns and structures of a musical piece. We build upon the recently proposed Masker-Denoiser (MaD) architecture and we enhance it with the Twin Networks, a technique to regularize a recurrent generative network using a backward running copy of the network. We evaluate our method using the Demixing Secret Dataset and we obtain an increment to signal-to-distortion ratio (SDR) of 0.37 dB and to signal-to-interference ratio (SIR) of 0.23 dB, compared to previous SOTA results. Konstantinos Drossos, Stylianos I. Mimilakis, Dmitriy Serdyuk, Gerald Schuller, Tuomas Virtanen, Yoshua Bengio |
IJCNN | 5 |
| 2018 | Reducing Interference with Phase Recovery in DNN-based Monaural Singing Voice SeparationabstractS.332-336 Paul Magron, Konstantinos Drossos, Stylianos I. Mimilakis, Tuomas Virtanen |
INTERSPEECH | 4 |
| 2018 | Expectation-Maximization Algorithms for Itakura-Saito Nonnegative Matrix FactorizationabstractInternational audience Paul Magron, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2018 | Multichannel Blind Sound Source Separation Using Spatial Covariance Model With Level and Time Differences and Nonnegative Matrix FactorizationabstractThis paper presents an algorithm for multichannel sound source separation using explicit modeling of level and time differences in source spatial covariance matrices (SCM). We propose a novel SCM model in which the spatial properties are modeled by the weighted sum of direction of arrival (DOA) kernels. DOA kernels are obtained as the combination of phase and level difference covariance matrices representing both time and level differences between microphones for a grid of predefined source directions. The proposed SCM model is combined with the NMF model for the magnitude spectrograms. Opposite to other SCM models in the literature, in this work, source localization is implicitly defined in the model and estimated during the signal factorization. Therefore, no localization preprocessing is required. Parameters are estimated using complex-valued nonnegative matrix factorization with both Euclidean distance and Itakura-Saito divergence. Separation performance of the proposed system is evaluated using the two-channel SiSEC development dataset and four channels signals recorded in a regular room with moderate reverberation. Finally, a comparison to other state-of-the-art methods is performed, showing better achieved separation performance in terms of SIR and perceptual measures. Julio J. Carabias-Orti, Joonas Nikunen, Tuomas Virtanen, Pedro Vera-Candeas |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Detection and Classification of Acoustic Scenes and Events: Outcome of the DCASE 2016 ChallengeabstractPublic evaluation campaigns and datasets promote active development in target research areas, allowing direct comparison of algorithms. The second edition of the challenge on detection and classification of acoustic scenes and events (DCASE 2016) has offered such an opportunity for development of the state-of-the-art methods, and succeeded in drawing together a large number of participants from academic and industrial backgrounds. In this paper, we report on the tasks and outcomes of the DCASE 2016 challenge. The challenge comprised four tasks: acoustic scene classification, sound event detection in synthetic audio, sound event detection in real-life audio, and domestic audio tagging. We present each task in detail and analyze the submitted systems in terms of design and performance. We observe the emergence of deep learning as the most popular classification method, replacing the traditional approaches based on Gaussian mixture models and support vector machines. By contrast, feature representations have not changed substantially throughout the years, as mel frequency-based representations predominate in all tasks. The datasets created for and used in DCASE 2016 are publicly available and are a valuable resource for further research. Annamaria Mesaros, Toni Heittola, Emmanouil Benetos, Peter Foster, Mathieu Lagrange, Tuomas Virtanen, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2018 | Separation of Moving Sound Sources Using Multichannel NMF and Acoustic TrackingabstractIn this paper, we propose a method for separation of moving sound sources. The method is based on first tracking the sources and then estimation of source spectrograms using multichannel nonnegative matrix factorization (NMF) and extracting the sources from the mixture by single-channel Wiener filtering. We propose a novel multichannel NMF model with time-varying mixing of the sources denoted by spatial covariance matrices (SCM) and provide update equations for optimizing model parameters minimizing squared Frobenius norm. The SCMs of the model are obtained based on estimated directions of arrival of tracked sources at each time frame. The evaluation is based on established objective separation criteria and using real recordings of two and three simultaneous moving sound sources. The compared methods include conventional beamforming and ideal ratio mask separation. The proposed method is shown to exceed the separation quality of other evaluated blind approaches according to all measured quantities. Additionally, we evaluate the method's susceptibility toward tracking errors by comparing the separation quality achieved using annotated ground truth source trajectories. Joonas Nikunen, Aleksandr Diment, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | ASR in Classroom Today: Automatic Visualization of Conceptual Network in Science Classrooms
Daniela Caballero, Roberto Araya, Hanna Kronholm, Jouni Viiri, André Mansikkaniemi, Sami Lehesvuori, Tuomas Virtanen, Mikko Kurimo |
EC-TEL | 7 |
| 2017 | Sound event detection using spatial features and convolutional recurrent neural networkabstractThis paper proposes to use low-level spatial features extracted from multichannel audio for sound event detection. We extend the convolutional recurrent neural network to handle more than one type of these multichannel features by learning from each of them separately in the initial stages. We show that instead of concatenating the features of each channel into a single feature vector the network learns sound events in multichannel audio better when they are presented as separate layers of a volume. Using the proposed spatial features over monaural features on the same network gives an absolute F-score improvement of 6.1% on the publicly available TUT-SED 2016 dataset and 2.7% on the TUT-SED 2009 dataset that is fifteen times larger. Sharath Adavanne, Pasi Pertilä, Tuomas Virtanen |
ICASSP | 3 |
| 2017 | Active learning for sound event classification by clustering unlabeled dataabstractThis paper proposes a novel active learning method to save annotation effort when preparing material to train sound event classifiers. K-medoids clustering is performed on unlabeled sound segments, and medoids of clusters are presented to annotators for labeling. The annotated label for a medoid is used to derive predicted labels for other cluster members. The obtained labels are used to build a classifier using supervised training. The accuracy of the resulted classifier is used to evaluate the performance of the proposed method. The evaluation made on a public environmental sound dataset shows that the proposed method outperforms reference methods (random sampling, certainty-based active learning and semi-supervised learning) with all simulated labeling budgets, the number of available labeling responses. Through all the experiments, the proposed method saves 50%-60% labeling budget to achieve the same accuracy, with respect to the best reference method. Shuyang Zhao, Toni Heittola, Tuomas Virtanen |
ICASSP | 3 |
| 2017 | A convolutional neural network approach for acoustic scene classificationabstractThis paper presents a novel application of convolutional neural networks (CNNs) for the task of acoustic scene classification (ASC). We here propose the use of a CNN trained to classify short sequences of audio, represented by their log-mel spectrogram. We also introduce a training method that can be used under particular circumstances in order to make full use of small datasets. The proposed system is tested and evaluated on three different ASC datasets and compared to other state-of-the-art systems which competed in the “Detection and Classification of Acoustic Scenes and Events” (DCASE) challenges held in 20161and 2013. The best accuracy scores obtained by our system on the DCASE 2016 datasets are 79.0% (development) and 86.2% (evaluation), which constitute a 6.4% and 9% improvements with respect to the baseline system. Finally, when tested on the DCASE 2013 evaluation dataset, the proposed system manages to reach a 77.0% accuracy, improving by 1% the challenge winner's score. Michele Valenti, Stefano Squartini, Aleksandr Diment, Giambattista Parascandolo, Tuomas Virtanen |
IJCNN | 5 |
| 2017 | Convolutional Recurrent Neural Networks for Polyphonic Sound Event DetectionabstractSound events often occur in unstructured environments where they exhibit wide variations in their frequency content and temporal structure. Convolutional neural networks (CNNs) are able to extract higher level features that are invariant to local spectral and temporal variations. Recurrent neural networks (RNNs) are powerful in learning the longer term temporal context in the audio signals. CNNs and RNNs as classifiers have recently shown improved performances over established methods in various sound recognition tasks. We combine these two approaches in a convolutional recurrent neural network (CRNN) and apply it on a polyphonic sound event detection task. We compare the performance of the proposed CRNN method with CNN, RNN, and other established methods, and observe a considerable improvement for four different datasets consisting of everyday sound events. Emre Çakir, Giambattista Parascandolo, Toni Heittola, Heikki Huttunen, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2017 | Binary Non-Negative Matrix Deconvolution for Audio Dictionary LearningabstractIn this study, we propose an unsupervised method for dictionary learning in audio signals. The new method, called binary nonnegative matrix deconvolution (BNMD), is developed and used to discover patterns from magnitude-scale spectrograms. The BNMD models an audio spectrogram as a sum of delayed patterns having binary gains (activations). Only small subsets of patterns can be active for a given spectrogram excerpt. The proposed method was applied to speaker identification and separation tasks. The experimental results show that dictionaries obtained by the BNMD bring much higher speaker identification accuracies averaged over a range of SNRs from -6 dB to 9 dB (91.3%) than the NMD-based dictionaries (37.8-75.4%). The BNMD also gives a benefit over dictionaries obtained using vector quantization (87.8%). For bigger dictionaries the difference between the BNMD and the vector quantization (VQ) is getting smaller. For the speech separation task the BNMD dictionary gave a slight improvement over the VQ. Szymon Drgas, Tuomas Virtanen, Jörg Lücke, Antti Hurmalainen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | Introduction to the Special Section on Sound Scene and Event AnalysisabstractThe papers in this special section are devoted to the growing field of acoustic scene classification and acoustic event recognition. Machine listening systems still have difficulties to reach the ability of human listeners in the analysis of realistic acoustic scenes. If sustained research efforts have been made for decades in speech recognition, speaker identification and to a lesser extent in music information retrieval, the analysis of other types of sounds, such as environmental sounds, is the subject of growing interest from the community and is targeting an ever increasing set of audio categories. This problem appears to be particularly challenging due to the large variety of potential sound sources in the scene, which may in addition have highly different acoustic characteristics, especially in bioacoustics. Furthermore, in realistic environments, multiple sources are often present simultaneously, and in reverberant conditions. Gaël Richard, Tuomas Virtanen, Juan Pablo Bello, Nobutaka Ono, Hervé Glotin |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Recurrent neural networks for polyphonic sound event detection in real life recordingsabstractIn this paper we present an approach to polyphonic sound event detection in real life recordings based on bi-directional long short term memory (BLSTM) recurrent neural networks (RNNs). A single multilabel BLSTM RNN is trained to map acoustic features of a mixture signal consisting of sounds from multiple classes, to binary activity indicators of each event class. Our method is tested on a large database of real-life recordings, with 61 classes (e.g. music, car, speech) from 10 different everyday contexts. The proposed method outperforms previous approaches by a large margin, and the results are further improved using data augmentation techniques. Overall, our system reports an average F1-score of 65.5% on 1 second blocks and 64.7% on single frames, a relative improvement over previous state-of-the-art approach of 6.8% and 15.1% respectively. Giambattista Parascandolo, Heikki Huttunen, Tuomas Virtanen |
ICASSP | 3 |
| 2016 | Filterbank learning for deep neural network based polyphonic sound event detectionabstractDeep learning techniques such as deep feedforward neural networks and deep convolutional neural networks have recently been shown to improve the performance in sound event detection compared to traditional methods such as Gaussian mixture models. One of the key factors of this improvement is the capability of deep architectures to automatically learn higher levels of acoustic features in each layer. In this work, we aim to combine the feature learning capabilities of deep architectures with the empirical knowledge of human perception. We use the first layer of a deep neural network to learn a mapping from a high-resolution magnitude spectrum to smaller amount of frequency bands, which effectively learns a filterbank for the sound event detection task. We initialize the first hidden layer weights to match with the perceptually motivated mel filterbank magnitude response. We also integrate this initialization scheme with context windowing by using an appropriately constrained deep convolutional neural network. The proposed method does not only result with better detection accuracy, but also provides insight on the frequencies deemed essential for better discrimination of given sound events. Emre Çakir, Ezgi C. Ozan, Tuomas Virtanen |
IJCNN | 3 |
| 2016 | Binaural rendering of microphone array captures based on source separation
Joonas Nikunen, Aleksandr Diment, Tuomas Virtanen, Miikka Vilermo |
Speech Commun. | 3 |
| 2016 | Blind Separation of Audio Mixtures Through Nonnegative Tensor Factorization of Modulation SpectrogramsabstractThis paper presents an algorithm for unsupervised single-channel source separation of audio mixtures. The approach specifically addresses the challenging case of separation where no training data are available. By representing mixtures in the modulation spectrogram (MS) domain, we exploit underlying similarities in patterns present across frequency. A three-dimensional tensor factorization is able to take advantage of these redundant patterns, and is used to separate a mixture into an approximated sum of components by minimizing a divergence cost. Furthermore, we show that the basic tensor factorization can be extended with convolution in time being used to improve separation results and provide update rules to learn components in such a manner. Following factorization, sources are reconstructed in the audio domain from estimated components using a novel approach based on reconstruction masks that are learned using MS activations, and then applied to a mixture spectrogram. We demonstrate that the proposed method produces superior separation performance to a spectrally based nonnegative matrix factorization approach, in terms of source-to-distortion ratio. We also compare separation with the perceptually motivated interference-related perceptual score metric and identify cases with higher performance. Tom Barker, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Exemplar-based speech enhancement for deep neural network based automatic speech recognitionabstractDeep neural network (DNN) based acoustic modelling has been successfully used for a variety of automatic speech recognition (ASR) tasks, thanks to its ability to learn higher-level information using multiple hidden layers. This paper investigates the recently proposed exemplar-based speech enhancement technique using coupled dictionaries as a pre-processing stage for DNN-based systems. In this setting, the noisy speech is decomposed as a weighted sum of atoms in an input dictionary containing exemplars sampled from a domain of choice, and the resulting weights are applied to a coupled output dictionary containing exemplars sampled in the short-time Fourier transform (STFT) domain to directly obtain the speech and noise estimates for speech enhancement. In this work, settings using input dictionary of exemplars sampled from the STFT, Mel-integrated magnitude STFT and modulation envelope spectra are evaluated. Experiments performed on the AURORA-4 database revealed that these pre-processing stages can improve the performance of the DNN-HMM-based ASR systems with both clean and multi-condition training. Deepak Baby, Jort F. Gemmeke, Tuomas Virtanen, Hugo Van hamme |
ICASSP | 3 |
| 2015 | Low-latency sound-source-separation using non-negative matrix factorisation with coupled analysis and synthesis dictionariesabstractFor real-time or close to real-time applications, sound source separation can be performed on-line, where new frames of incoming data for a mixture signal are processed as they arrive, at very low delay. We propose an approach which generates the separation filters for short synthesis frames to achieve low latency source separation, based on a compositional model mixture of the audio to be separated. Filter parameters are derived from a longer temporal context than the current processing frame through use of a longer analysis frame. A pair of dictionaries are used, one for analysis and one for reconstruction. With this approach we are able to increase separation performance at low latencies whilst retaining the low-latency provided by the use of short synthesis frames. The proposed data handling scheme and parameters can be adjusted to achieve real-time performance, given sufficient computational power. Low-latency output allows a human listener to use the results of such a separation scheme directly, without a perceptible delay. With the proposed method, separated source-to-distortion ratios (SDRs) can be improved by over 1 dB for latencies below 20 ms, without any affect on latency. Tom Barker, Tuomas Virtanen, Niels Henrik Pontoppidan |
ICASSP | 2 |
| 2015 | Similarity induced group sparsity for non-negative matrix factorisationabstractNon-negative matrix factorisations are used in several branches of signal processing and data analysis for separation and classification. Sparsity constraints are commonly set on the model to promote discovery of a small number of dominant patterns. In group sparse models, atoms considered to belong to a consistent group are permitted to activate together, while activations across groups are suppressed, reducing the number of simultaneously active sources or other structures. Whereas most group sparse models require explicit division of atoms into separate groups without addressing their mutual relations, we propose a constraint that permits dynamic relationships between atoms or groups, based on any defined distance measure. The resulting solutions promote approximation with components considered similar to each other. Evaluation results are shown for speech enhancement and noise robust speech and speaker recognition. Antti Hurmalainen, Rahim Saeidi, Tuomas Virtanen |
ICASSP | 3 |
| 2015 | Sound event detection in real life recordings using coupled matrix factorization of spectral representations and class activity annotationsabstractMethods for detection of overlapping sound events in audio involve matrix factorization approaches, often assigning separated components to event classes. We present a method that bypasses the supervised construction of class models. The method learns the components as a non-negative dictionary in a coupled matrix factorization problem, where the spectral representation and the class activity annotation of the audio signal share the activation matrix. In testing, the dictionaries are used to estimate directly the class activations. For dealing with large amount of training data, two methods are proposed for reducing the size of the dictionary. The methods were tested on a database of real life recordings, and outperformed previous approaches by over 10%. Annamaria Mesaros, Toni Heittola, Onur Dikmen, Tuomas Virtanen |
ICASSP | 4 |
| 2015 | Polyphonic sound event detection using multi label deep neural networksabstractIn this paper, the use of multi label neural networks are proposed for detection of temporally overlapping sound events in realistic environments. Real-life sound recordings typically have many overlapping sound events, making it hard to recognize each event with the standard sound event detection methods. Frame-wise spectral-domain features are used as inputs to train a deep neural network for multi label classification in this work. The model is evaluated with recordings from realistic everyday environments and the obtained overall accuracy is 63.8%. The method is compared against a state-of-the-art method using non-negative matrix factorization as a pre-processing stage and hidden Markov models as a classifier. The proposed method improves the accuracy by 19% percentage points overall. Emre Çakir, Toni Heittola, Heikki Huttunen, Tuomas Virtanen |
IJCNN | 4 |
| 2015 | Noise robust speaker recognition with convolutive sparse codingabstractRecognition and classification of speech content in everyday environments is challenging due to the large diversity of real-world noise sources, which may also include competing speech. At signal-to-noise ratios below 0 dB, a majority of features may become corrupted, severely degrading the performance of clas-sifiers built upon clean observations of a target class. As the energy and complexity of competing sources increase, their ex-plicit modelling becomes integral for successful detection and classification of target speech. We have previously demon-strated how non-negative compositional modelling in a spec-trogram space is suitable for robust recognition of speech and speakers even at low SNRs. In this work, the sparse coding approach is extended to cover the whole separation and clas-sification chain to recognise the speaker of short utterances in difficult noise environments. A convolutive matrix factorisation and coding system is evaluated on 2nd CHiME Track 1 data. Over 98 % average speaker recognition accuracy is achieved for shorter than three second utterances at +9...-6 dB SNR, illus-trating the system’s performance in challenging conditions. Index Terms: speaker recognition, noise robustness, composi-tional models, sparse coding, non-negative matrix factorization Antti Hurmalainen, Rahim Saeidi, Tuomas Virtanen |
INTERSPEECH | 3 |
| 2015 | Coupled Dictionaries for Exemplar-Based Speech Enhancement and Automatic Speech RecognitionabstractExemplar-based speech enhancement systems work by decomposing the noisy speech as a weighted sum of speech and noise exemplars stored in a dictionary and use the resulting speech and noise estimates to obtain a time-varying filter in the full-resolution frequency domain to enhance the noisy speech. To obtain the decomposition, exemplars sampled in lower dimensional spaces are preferred over the full-resolution frequency domain for their reduced computational complexity and the ability to better generalize to unseen cases. But the resulting filter may be sub-optimal as the mapping of the obtained speech and noise estimates to the full-resolution frequency domain yields a low-rank approximation. This paper proposes an efficient way to directly compute the full-resolution frequency estimates of speech and noise using coupled dictionaries: an input dictionary containing atoms from the desired exemplar space to obtain the decomposition and a coupled output dictionary containing exemplars from the full-resolution frequency domain. We also introduce modulation spectrogram features for the exemplar-based tasks using this approach. The proposed system was evaluated for various choices of input exemplars and yielded improved speech enhancement performances on the AURORA-2 and AURORA-4 databases. We further show that the proposed approach also results in improved word error rates (WERs) for the speech recognition tasks using HMM-GMM and deep-neural network (DNN) based systems. Deepak Baby, Tuomas Virtanen, Jort F. Gemmeke, Hugo Van hamme |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Coupled dictionary training for exemplar-based speech enhancementabstractIn exemplar-based speech enhancement systems, lower dimensional features are preferred over the full-scale DFT features for their reduced computational complexity and the ability to better generalize for the unseen cases. But in order to obtain the Wiener-like filter for noisy DFT enhancement, the speech and noise estimates obtained in the feature space need to be mapped to the DFT space, which yield a low-rank approximation of the estimates resulting in a sub-optimal filter. This paper proposes a novel method using coupled dictionaries where the exemplars for the required feature space and the DFT space are jointly extracted and the estimates are directly obtained in the DFT space following the decomposition in the chosen feature space. Simulation experiments revealed that the proposed approach, where the activations of exemplars calculated using the Mel resolution are directly used to obtain the Wiener filter in the DFT space, results in improved signal-to-distortion ratio (SDR) when compared to the system without coupled dictionaries. To further motivate the use of coupled dictionaries, the paper also investigates the use of modulation envelope features for the exemplar-based speech enhancement. Deepak Baby, Tuomas Virtanen, Tom Barker, Hugo Van hamme |
ICASSP | 2 |
| 2014 | Ultrasound-coupled semi-supervised nonnegative matrix factorisation for speech enhancementabstractWe present an extension to an existing speech enhancement technique, whereby the incorporation of easily obtained Doppler-based ultrasound data, obtained from frequency shifts caused by a talker's mouth movements, is shown to improve speech enhancement results. Noisy speech mixtures were enhanced using semi-supervised nonnegative matrix factorisation (NMF). Ultrasound data recorded alongside the speech is transformed into the spectral domain and used additionally to audio in the mixture to be separated. Speech components are learned from a training set, whilst noise components are estimated from the mixture signal. We show that the ultrasound data can improve source-to-distortion ratios for the enhanced speech, relative to both the non-ultrasound NMF case and an established Wiener filter-based speech enhancement method. Tom Barker, Tuomas Virtanen, Olivier Delhomme |
ICASSP | 2 |
| 2014 | Multichannel audio separation by direction of arrival based spatial covariance model and non-negative matrix factorizationabstractThis paper studies multichannel audio separation using non-negative matrix factorization (NMF) combined with a new model for spatial covariance matrices (SCM). The proposed model for SCMs is parameterized by source direction of arrival (DoA) and its parameters can be optimized to yield a spatially coherent solution over frequencies thus avoiding permutation ambiguity and spatial aliasing. The model constrains the estimation of SCMs to a set of geometrically possible solutions. Additionally we present a method for using a priori DoA information of the sources extracted blindly from the mixture for the initialization of the parameters of the proposed model. The simulations show that the proposed algorithm exceeds the separation quality of existing spatial separation methods. Joonas Nikunen, Tuomas Virtanen |
ICASSP | 2 |
| 2014 | Active-set newton algorithm for non-negative sparse coding of audioabstractWe propose a new algorithm to efficiently obtain non-negative sparse representations for audio. The spectrum of an audio signal is represented as a sparse linear combination of atoms taken from an overcomplete dictionary. The algorithm is based on minimizing the generalized Kullback-Leibler divergence between an observed magnitude spectrum and a non-negative linear combination of atoms, plus an ℓ1regularization term. The proposed method consists of an active-set method that iteratively updates a set of active atoms that have non-zero weights, using a Newton step where the weights of the active atoms are updated. The proposed method was evaluated using mixtures of two speakers, and it was shown to yield more than 10 times faster convergence in comparison to an established algorithm based on multiplicative update rules. Moreover, the ℓ1regularization was found to decrease the computation time and to improve the source separation performance. Tuomas Virtanen, Bhiksha Raj, Jort F. Gemmeke, Hugo Van hamme |
ICASSP | 1 |
| 2014 | Semi-supervised non-negative tensor factorisation of modulation spectrograms for monaural speech separationabstractThis paper details the use of a semi-supervised approach to audio source separation. Where only a single source model is available, the model for an unknown source must be estimated. A mixture signal is separated through factorisation of a feature-tensor representation, based on the modulation spectrogram. Harmonically related components tend to modulate in a similar fashion, and this redundancy of patterns can be isolated. This feature representation requires fewer parameters than spectrally based methods and so minimises overfitting. Following the tensor factorisation, the separated signals are reconstructed by learning appropriate Wiener-filter spectral parameters which have been constrained by activation parameters learned in the first stage. Strong results were obtained for two-speaker mixtures where source separation performance exceeded those used as benchmarks. Specifically, the proposed semi-supervised method outperformed both semi-supervised non-negative matrix factorisation and blind non-negative modulation spectrum tensor factorisation. Tom Barker, Tuomas Virtanen |
IJCNN | 2 |
| 2014 | Modelling primitive streaming of simple tone sequences through factorisation of modulation pattern tensorsabstractCopyright © 2014 ISCA. We present a novel method for determining how the perceptual organisation of simple alternating tone sequences is likely to occur in human listeners. By training a tensor model representation using features which incorporate both low-frequency modulation rate and phase, a set of components is learned. Test patterns are modelled using these learned components, and the sum of component activations is used to predict either an 'integrated' or 'segregated' auditory stream percept. We find that for the basic streaming paradigm tested, our proposed model and method is able to correctly predict either segregation or integration in the majority of cases. Tom Barker, Hugo Van hamme, Tuomas Virtanen |
INTERSPEECH | 3 |
| 2014 | Exemplar-based noise robust automatic speech recognition using modulation spectrogram featuresabstractWe propose a novel exemplar-based feature enhancement method for automatic speech recognition which uses coupled dictionaries: an input dictionary containing atoms sampled in the modulation (envelope) spectrogram domain and an output dictionary with atoms in the Mel or full-resolution frequency domain. The input modulation representation is chosen for its separation properties of speech and noise and for its relation with human auditory processing. The output representation is one which can be processed by the ASR back-end. The proposed method was investigated on the AURORA-2 and AURORA-4 databases and improved word error rates (WER) were obtained when compared to the system which uses Mel features in the input exemplars. The paper also proposes a hybrid system which combines the baseline and the proposed algorithm on the AURORA-2 database which in turn also yielded improvement over both the algorithms. Deepak Baby, Tuomas Virtanen, Jort F. Gemmeke, Tom Barker, Hugo Van hamme |
SLT | 2 |
| 2014 | Direction of Arrival Based Spatial Covariance Model for Blind Sound Source SeparationabstractThis paper addresses the problem of sound source separation from a multichannel microphone array capture via estimation of source spatial covariance matrix (SCM) of a short-time Fourier transformed mixture signal. In many conventional audio separation algorithms the source mixing parameter estimation is done separately for each frequency thus making them prone to errors and leading to suboptimal source estimates. In this paper we propose a SCM model which consists of a weighted sum of direction of arrival (DoA) kernels and estimate only the weights dependent on the source directions. In the proposed algorithm, the spatial properties of the sources become jointly optimized over all frequencies, leading to more coherent source estimates and mitigating the effect of spatial aliasing at high frequencies. The proposed SCM model is combined with a linear model for magnitudes and the parameter estimation is formulated in a complex-valued non-negative matrix factorization (CNMF) framework. Simulations consist of recordings done with a hand-held device sized array having multiple microphones embedded inside the device casing. Separation quality of the proposed algorithm is shown to exceed the performance of existing state of the art separation methods with two sources when evaluated by objective separation quality metrics. Joonas Nikunen, Tuomas Virtanen |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Exemplar-Based Sparse Representation With Residual Compensation for Voice ConversionabstractWe propose a nonparametric framework for voice conversion, that is, exemplar-based sparse representation with residual compensation. In this framework, a spectrogram is reconstructed as a weighted linear combination of speech segments, called exemplars, which span multiple consecutive frames. The linear combination weights are constrained to be sparse to avoid over-smoothing, and high-resolution spectra are employed in the exemplars directly without dimensionality reduction to maintain spectral details. In addition, a spectral compression factor and a residual compensation technique are included in the framework to enhance the conversion performances. We conducted experiments on the VOICES database to compare the proposed method with a large set of state-of-the-art baseline methods, including the maximum likelihood Gaussian mixture model (ML-GMM) with dynamic feature constraint and the partial least squares (PLS) regression based methods. The experimental results show that the objective spectral distortion of ML-GMM is reduced from 5.19 dB to 4.92 dB, and both the subjective mean opinion score and the speaker identification rate are increased from 2.49 and 73.50% to 3.15 and 79.50%, respectively, by the proposed method. The results also show the superiority of our method over PLS-based methods. In addition, the subjective listening tests indicate that the naturalness of the converted speech by our proposed method is comparable with that by the ML-GMM method with global variance constraint. Zhizheng Wu 0001, Tuomas Virtanen, Chng Eng Siong, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Learning state labels for sparse classification of speech with matrix deconvolutionabstractNon-negative spectral factorisation with long temporal context has been successfully used for noise robust recognition of speech in multi-source environments. Sparse classification from activations of speech atoms can be employed instead of conventional GMMs to determine speech state likelihoods. For accurate classification, correct linguistic state labels must be assigned to speech atoms. We propose using non-negative matrix deconvolution for learning the labels with algorithms closely matching a framework that separates speech from additive noises. Experiments on the 1st CHiME Challenge corpus show improvement in recognition accuracy over labels acquired from original atom sources or previously used least squares regression. The new approach also circumvents numerical issues encountered in previous learning methods, and opens up possibilities for new speech basis generation algorithms. Antti Hurmalainen, Tuomas Virtanen |
ASRU | 2 |
| 2013 | Exemplar-based joint channel and noise compensationabstractIn this paper two models for channel estimation in exemplar-based noise robust speech recognition are proposed. Building on a compositional model that models noisy speech and a combination of noise and speech atoms, the first model iteratively estimates a filter to best compensate the mismatch with the observed noisy speech. The second model estimates separate filters for the noise and speech atoms. We show that both models enable noise-robust ASR even if the channel characteristics of the noisy speech do not match those of the exemplars in the dictionary. Moreover, the second model, which is able to estimate separate filters for speech and noise, is shown to be robust even in the presence of bandwidth-limited sources. Jort F. Gemmeke, Tuomas Virtanen, Kris Demuynck |
ICASSP | 2 |
| 2013 | Supervised model training for overlapping sound events based on unsupervised source separationabstractSound event detection is addressed in the presence of overlapping sounds. Unsupervised sound source separation into streams is used as a preprocessing step to minimize the interference of overlapping events. This poses a problem in supervised model training, since there is no knowledge about which separated stream contains the targeted sound source. We propose two iterative approaches based on EM algorithm to select the most likely stream to contain the target sound: one by selecting always the most likely stream and another one by gradually eliminating the most unlikely streams from the training. The approaches were evaluated with a database containing recordings from various contexts, against the baseline system trained without applying stream selection. Both proposed approaches were found to give a reasonable increase of 8 percentage units in the detection accuracy. Toni Heittola, Annamaria Mesaros, Tuomas Virtanen, Moncef Gabbouj |
ICASSP | 3 |
| 2013 | Non-negative tensor factorisation of modulation spectrograms for monaural sound source separationabstractThis paper proposes an algorithm for separating monaural audio signals by non-negative tensor factorisation of modulation spectrograms. The modulation spectrogram is able to represent redundant patterns across frequency with similar features, and the tensor factorisation is able to isolate these patterns in an unsupervised way. The method overcomes the limitation of conventional non-negative matrix factorisation algorithms to utilise the redundancy of sounds in frequency. In the proposed method, separated sounds are synthesised by filtering the mixture signal with a Wiener-like filter generated from the estimated tensor factors. The proposed method was compared to conventional algorithms in unsupervised separation of mixtures of speech and music. Improved signal to distortion ratios were obtained compared to standard non-negative matrix factorisation and non-negative matrix deconvolution. Tom Barker, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2013 | Exemplar-based unit selection for voice conversion utilizing temporal informationabstractAlthough temporal information of speech has been shown to play an important role in perception, most of the voice conver-sion approaches assume the speech frames are independent of each other, thereby ignoring the temporal information. In this study, we improve conventional unit selection approach by us-ing exemplars which span multiple frames as base units, and also take temporal information constraint into voice conver-sion by using overlapping frames to generate speech parame-ters. This approach thus provides more stable concatenation cost and avoids discontinuity problem in conventional unit se-lection approach. The proposed method also keeps away from the over-smoothing problem in the mainstream joint density Gaussian mixture model (JD-GMM) based conversion method by directly using target speaker’s training data for synthesizing the converted speech. Both objective and subjective evaluations indicate that our proposed method outperforms JD-GMM and conventional unit selection methods. Index Terms: Voice conversion, unit selection, multi-frame ex-emplar, temporal information Zhizheng Wu 0001, Tuomas Virtanen, Tomi Kinnunen, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2013 | Modelling non-stationary noise with spectral factorisation in automatic speech recognition
Antti Hurmalainen, Jort F. Gemmeke, Tuomas Virtanen |
Comput. Speech Lang. | 3 |
| 2013 | On the human ability to discriminate audio ambiances from similar locations of an urban environment
Dani Korpi, Toni Heittola, Timo Partala, Antti J. Eronen, Annamaria Mesaros, Tuomas Virtanen |
Pers. Ubiquitous Comput. | 6 |
| 2013 | Active-Set Newton Algorithm for Overcomplete Non-Negative Representations of AudioabstractThis paper proposes a computationally efficient algorithm for estimating the non-negative weights of linear combinations of the atoms of large-scale audio dictionaries, so that the generalized Kullback-Leibler divergence between an audio observation and the model is minimized. This linear model has been found useful in many audio signal processing tasks, but the existing algorithms are computationally slow when a large number of atoms is used. The proposed algorithm is based on iteratively updating a set of active atoms, with the weights updated using the Newton method and the step size estimated such that the weights remain non-negative. Algorithm convergence evaluations on representing audio spectra that are mixtures of two speakers show that with all the tested dictionary sizes the proposed method reaches a much lower value of the divergence than can be obtained by conventional algorithms, and is up to 8 times faster. A source separation evaluation revealed that when using large dictionaries, the proposed method produces a better separation quality in less time. Tuomas Virtanen, Jort F. Gemmeke, Bhiksha Raj |
IEEE Trans. Speech Audio Process. | 1 |
| 2012 | Modelling spectro-temporal dynamics in factorisation-based noise-robust automatic speech recognitionabstractNon-negative spectral factorisation has been used successfully for separation of speech and noise in automatic speech recognition, both in feature-enhancing front-ends and in direct classification. In this work, we propose employing spectro-temporal 2D filters to model dynamic properties of Mel-scale spectrogram patterns in addition to static magnitude features. The results are evaluated using an exemplar-based sparse classifier on the CHiME noisy speech database. After optimisation of static features and modelling of temporal dynamics with derivative features, we achieve 87.4% average score over SNRs from 9 to -6 dB, reducing the word error rate by 28.1% from our previous static-only features. Antti Hurmalainen, Tuomas Virtanen |
ICASSP | 2 |
| 2012 | Non-negative matrix factorization for highly noise-robust ASR: To enhance or to recognize?abstractThis paper proposes a multi-stream speech recognition system that combines information from three complementary analysis methods in order to improve automatic speech recognition in highly noisy and reverberant environments, as featured in the 2011 PASCAL CHiME Challenge. We integrate word predictions by a bidirectional Long Short-Term Memory recurrent neural network and non-negative sparse classification (NSC) into a multi-stream Hidden Markov Model using convolutive non-negative matrix factorization (NMF) for speech enhancement. Our results suggest that NMF-based enhancement and NSC are complementary despite their overlap in methodology, reaching up to 91.9% average keyword accuracy on the Challenge test set at signal-to-noise ratios from -6 to 9 dB-the best result reported so far on these data. Felix Weninger, Martin Wöllmer, Jürgen T. Geiger, Björn W. Schuller, Jort F. Gemmeke, Antti Hurmalainen, Tuomas Virtanen, Gerhard Rigoll |
ICASSP | 7 |
| 2012 | Group Sparsity for Speaker Identity Discrimination in Factorisation-based Speech RecognitionabstractSpectrogram factorisation using a dictionary of spectro-temporal atoms has been successfully employed to separate a mixed audio signal into its source components. When atoms from multiple sources are included in a combined dictionary, the relative weights of activated atoms reveal likely sources as well as the content of each source. Enforcing sparsity on the activation weights produces solutions, where only a small number of atoms are active at a time. In this paper we pro-pose using group sparsity to restrict simultaneous activation of sources, allowing us to discover the identity of an unknown speaker from multiple candidates, and further to recognise the phonetic content more reliably with a narrowed down subset of atoms belonging to the most likely speakers. An evalua-tion on the CHiME corpus shows that the use of group sparsity improves the results of noise robust speaker identification and speech recognition using speaker-dependent models. Index Terms: group sparsity, speech recognition, speaker iden-tification, spectrogram factorization Antti Hurmalainen, Rahim Saeidi, Tuomas Virtanen |
INTERSPEECH | 3 |
| 2012 | Voice Conversion Using Dynamic Kernel Partial Least Squares RegressionabstractA drawback of many voice conversion algorithms is that they rely on linear models and/or require a lot of tuning. In addition, many of them ignore the inherent time-dependency between speech features. To address these issues, we propose to use dynamic kernel partial least squares (DKPLS) technique to model nonlinearities as well as to capture the dynamics in the data. The method is based on a kernel transformation of the source features to allow non-linear modeling and concatenation of previous and next frames to model the dynamics. Partial least squares regression is used to find a conversion function that does not overfit to the data. The resulting DKPLS algorithm is a simple and efficient algorithm and does not require massive tuning. Existing statistical methods proposed for voice conversion are able to produce good similarity between the original and the converted target voices but the quality is usually degraded. The experiments conducted on a variety of conversion pairs show that DKPLS, being a statistical method, enables successful identity conversion while achieving a major improvement in the quality scores compared to the state-of-the-art Gaussian mixture-based model. In addition to enabling better spectral feature transformation, quality is further improved when aperiodicity and binary voicing values are converted using DKPLS with auxiliary information from spectral features. Elina Helander, Hanna Silén, Tuomas Virtanen, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 3 |
| 2011 | Non-negative matrix deconvolution in noise robust speech recognitionabstractHigh noise robustness has been achieved in speech recognition by using sparse exemplar-based methods with spectrogram windows spanning up to 300 ms. A downside is that a large exemplar dictionary is required to cover sufficiently many spectral patterns and their temporal alignments within windows. We propose a recognition system based on a shift-invariant convolutive model, where exemplar activations at all the possible temporal positions jointly reconstruct an utterance. Recognition rates are evaluated using the AURORA-2 database, containing spoken digits with noise ranging from clean speech to -5 dB SNR. We obtain results superior to those, where the activations were found independently for each overlapping window. Antti Hurmalainen, Jort F. Gemmeke, Tuomas Virtanen |
ICASSP | 3 |
| 2011 | Uncertainty Measures for Improving Exemplar-Based Source SeparationabstractThis work studies the use of observation uncertainty mea-sures for improving the speech recognition performance of an exemplar-based source separation based front end. To generate the observation uncertainty estimates for the enhanced features, we propose the use of heuristic methods based on the sparse representation of the noisy signal in the exemplar-based source separation algorithm. The effectiveness of the proposed mea-sures is evaluated in a large vocabulary noisy speech recogni-tion task. The best proposed measure achieved relative error reductions up to 18 % over the baseline feature enhancement method without uncertainty measures. Index Terms: robustness, speech recognition, source separa-tion, observation uncertainties Heikki Kallasjoki, Ulpu Remes, Jort F. Gemmeke, Tuomas Virtanen, Kalle J. Palomäki |
INTERSPEECH | 4 |
| 2011 | Mapping Sparse Representation to State Likelihoods in Noise-Robust Automatic Speech Recognitionabstractstatus: Published Katariina Mahkonen, Antti Hurmalainen, Tuomas Virtanen, Jort F. Gemmeke |
INTERSPEECH | 3 |
| 2011 | Phoneme-Dependent NMF for Speech Enhancement in Monaural MixturesabstractThe problem of separating speech signals out of monaural mix-tures (with other non-speech or speech signals) has become in-creasingly popular in recent times. Among the various solutions proposed, the most popular methods are based on compositional models such as non-negative matrix factorization (NMF) and latent variable models. Although these techniques are highly effective they largely ignore the inherently phonetic nature of speech. In this paper we present a phoneme-dependent NMF-based algorithm to separate speech from monaural mixtures. Experiments performed on speech mixed with music indicate that the proposed algorithm can result in significant improve-ment in separation performance, over conventional NMF-based separation. Index terms: Monaural signal separation, speech enhance-ment, restoration, Non-negative matrix factorization. 1. Bhiksha Raj, Rita Singh, Tuomas Virtanen |
INTERSPEECH | 3 |
| 2011 | Exemplar-Based Sparse Representations for Noise Robust Automatic Speech RecognitionabstractThis paper proposes to use exemplar-based sparse representations for noise robust automatic speech recognition. First, we describe how speech can be modeled as a linear combination of a small number of exemplars from a large speech exemplar dictionary. The exemplars are time-frequency patches of real speech, each spanning multiple time frames. We then propose to model speech corrupted by additive noise as a linear combination of noise and speech exemplars, and we derive an algorithm for recovering this sparse linear combination of exemplars from the observed noisy speech. We describe how the framework can be used for doing hybrid exemplar-based/HMM recognition by using the exemplar-activations together with the phonetic information associated with the exemplars. As an alternative to hybrid recognition, the framework also allows us to take a source separation approach which enables exemplar-based feature enhancement as well as missing data mask estimation. We evaluate the performance of these exemplar-based methods in connected digit recognition on the AURORA-2 database. Our results show that the hybrid system performed substantially better than source separation or missing data mask estimation at lower signal-to-noise ratios (SNRs), achieving up to 57.1% accuracy at SNR = -5 dB. Although not as effective as two baseline recognizers at higher SNRs, the novel approach offers a promising direction of future research on exemplar-based ASR. Jort F. Gemmeke, Tuomas Virtanen, Antti Hurmalainen |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Noise robust exemplar-based connected digit recognitionabstractThis paper proposes a noise robust exemplar-based speech recognition system where noisy speech is modeled as a linear combination of a set of speech and noise exemplars. The method works by finding a small number of labeled exemplars in a very large collection of speech and noise exemplars that jointly approximate the observed speech signal. We represent the exemplars using mel-energies, which allows modeling the summation of speech and noise, and estimate the activations of the exemplars by minimizing the generalized Kullback-Leibler divergence between the observations and the model. The activations of the speech exemplars are directly being used for recognition. This approach proves to be promising, achieving up to 55.8% accuracy at signal-to-noise ratio -5 dB on the AURORA-2 connected digit recognition task. Jort F. Gemmeke, Tuomas Virtanen |
ICASSP | 2 |
| 2010 | Sound source separation in monaural music signals using excitation-filter model and em algorithmabstractThis paper proposes a method for separating the signals of individual musical instruments from monaural musical audio. The mixture signal is modeled as a sum of the spectra of individual musical sounds which are further represented as a product of excitations and filters. The excitations are restricted to harmonic spectra and their fundamental frequencies are estimated in advance using a multipitch estimator, whereas the filters are restricted to have smooth frequency responses by modeling them as a sum of elementary functions on Mel-frequency scale. A novel expectation-maximization (EM) algorithm is proposed which jointly learns the filter responses and organizes the excitations (musical notes) to filters (instruments). In simulations, the method achieved over 5 dB SNR improvement compared to the mixture signals when separating two or three musical instruments from each other. A slight further improvement was achieved by utilizing musical properties in the initialization of the algorithm. Anssi Klapuri, Tuomas Virtanen, Toni Heittola |
ICASSP | 2 |
| 2010 | Recognition of phonemes and words in singingabstractThis paper studies the influence of n-gram language models in the recognition of sung phonemes and words. We train uni-, bi-, and trigram language models for phonemes and bi- and trigrams for words. The word-level language model is estimated from a textual lyrics database. In the recognition we use a hidden Markov model based phonetic recognizer adapted to singing voice. The models were tested on monophonic singing and on vocal lines separated from polyphonic music. On clean singing the phoneme recognition accuracies varied from 20% (no language model) to 39% (bigram) and on polyphonic music from 6% (no language model) to 20% (bigram). In word recognition, one fifth of the words were recognized in clean singing, the performance being lower on polyphonic music. We study the use of the recognition results in a query-by-singing application. Using the recognized words, we retrieve the songs by searching for the text in a text lyrics database. For the word recognition system having only 24% correct recognition rate, the first retrieved song is correct in 57% of the test cases. Annamaria Mesaros, Tuomas Virtanen |
ICASSP | 2 |
| 2010 | Noise-to-mask ratio minimization by weighted non-negative matrix factorizationabstractThis paper proposes a novel algorithm for minimizing the perceptual distortion in non-negative matrix factorization (NMF) based audio representation. We formulate the noise-to-mask ratio audio quality criterion in a form where it can be used in NMF and propose an algorithm for optimizing the criterion. We also propose a method for compensating the spreading of the representation error in the synthesis filterbank. The objective perceptual quality produced by the proposed method is found to outperform all the reference methods. We also study the trade-off between the window length and the rank of factorization with a fixed data rate, and find that the best performance is obtained with window lengths between 10 and 30 ms. Joonas Nikunen, Tuomas Virtanen |
ICASSP | 2 |
| 2010 | Artificial and online acquired noise dictionaries for noise robust ASRabstractContains fulltext : 86540.pdf (Publisher’s version ) (Open Access) Jort F. Gemmeke, Tuomas Virtanen |
INTERSPEECH | 2 |
| 2010 | Non-negative matrix factorization based compensation of music for automatic speech recognitionabstractThis paper proposes to use non-negative matrix factorization based speech enhancement in robust automatic recognition of mixtures of speech and music. We represent magnitude spectra of noisy speech signals as the non-negative weighted linear combination of speech and noise spectral basis vectors, that are obtained from training corpora of speech and music. We use overcomplete dic-tionaries consisting of random exemplars of the training data. The method is tested on theWall Street Journal large vocabulary speech corpus which is artificially corrupted with polyphonic music from the RWC music database. Various music styles and speech-to-music ratios are evaluated. The proposed methods are shown to produce a consistent, significant improvement on the recog-nition performance in the comparison with the baseline method. Audio demonstrations of the enhanced signals are available at Bhiksha Raj, Tuomas Virtanen, Sourish Chaudhuri, Rita Singh |
INTERSPEECH | 2 |
| 2010 | State-based labelling for a sparse representation of speech and its application to robust speech recognitionabstractContains fulltext : 86590.pdf (Publisher’s version ) (Open Access) Tuomas Virtanen, Jort F. Gemmeke, Antti Hurmalainen |
INTERSPEECH | 1 |
| 2010 | Voice Conversion Using Partial Least Squares RegressionabstractVoice conversion can be formulated as finding a mapping function which transforms the features of the source speaker to those of the target speaker. Gaussian mixture model (GMM)-based conversion is commonly used, but it is subject to overfitting. In this paper, we propose to use partial least squares (PLS)-based transforms in voice conversion. To prevent overfitting, the degrees of freedom in the mapping can be controlled by choosing a suitable number of components. We propose a technique to combine PLS with GMMs, enabling the use of multiple local linear mappings. To further improve the perceptual quality of the mapping where rapid transitions between GMM components produce audible artefacts, we propose to low-pass filter the component posterior probabilities. The conducted experiments show that the proposed technique results in better subjective and objective quality than the baseline joint density GMM approach. In speech quality conversion preference tests, the proposed method achieved 67% preference score against the smoothed joint density GMM method and 84% preference score against the unsmoothed joint density GMM method. In objective tests the proposed method produced a lower Mel-cepstral distortion than the reference methods. Elina Helander, Tuomas Virtanen, Jani Nurminen, Moncef Gabbouj |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Representing Musical Sounds With an Interpolating State ModelabstractA computationally efficient algorithm is proposed for modeling and representing time-varying musical sounds. The aim is to encode individual sounds and not the statistical properties of several sounds representing a certain class. A given sequence of acoustic feature vectors is modeled by finding such a set of ¿states¿ (anchor points in the feature space) that the input data can be efficiently represented by interpolating between them. The proposed interpolating state model is generic and can be used to represent any multidimensional data sequence. In this paper, it is applied to represent musical instrument sounds in a compact and accurate form. Simulation experiments were carried out which show that the proposed method clearly outperforms the conventional vector quantization approach where the acoustic feature data is k-means clustered and the feature vectors are replaced by the corresponding cluster centroids. The computational complexity of the proposed algorithm as a function of the input sequence length T is O(T log T). Anssi Klapuri, Tuomas Virtanen |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Interpolating hidden Markov model and its application to automatic instrument recognitionabstractThis paper proposes an interpolating extension to hidden Markov models (HMMs), which allows more accurate modeling of natural sounds sources. The model is able to produce observations from distributions which are interpolated between discrete HMM states. The model uses Gaussian mixture state emission densities, and the interpolation is implemented by introducing interpolating states in which the mixture weights, means, and variances are interpolated from the discrete HMM state densities. We propose an algorithm extended from the Baum-Welch algorithm for estimating the parameters of the interpolating model. The model was evaluated in automatic instrument classification task, where it produced systematically better recognition accuracy than a baseline HMM recognition algorithm. Tuomas Virtanen, Toni Heittola |
ICASSP | 1 |
| 2008 | Bayesian extensions to non-negative matrix factorisation for audio signal modellingabstractWe describe the underlying probabilistic generative signal model of non-negative matrix factorisation (NMF) and propose a realistic conjugate priors on the matrices to be estimated. A conjugate Gamma chain prior enables modelling the spectral smoothness of natural sounds in general, and other prior knowledge about the spectra of the sounds can be used without resorting to too restrictive techniques where some of the parameters are fixed. The resulting algorithm, while retaining the attractive features of standard NMF such as fast convergence and easy implementation, outperforms existing NMF strategies in a single channel audio source separation and detection task. Tuomas Virtanen, A. Taylan Cemgil, Simon J. Godsill |
ICASSP | 1 |
| 2008 | Accompaniment separation and karaoke application based on automatic melody transcriptionabstractWe propose a method for separating accompaniment from polyphonic music and its karaoke application, both based on automatic melody transcription. First, the method transcribes the lead-vocal melody of an existing polyphonic music piece, where the transcription consists of a MIDI note sequence and a detailed fundamental frequency (F0) trajectory for each note. Based on the note F0 trajectories, the method uses sinusoidal modeling to estimate, synthesize, and remove the lead vocals in the piece, thus producing separated accompaniment of the piece. User sings along with the separated accompaniment similar to karaoke while the user singing can be tuned to the transcribed melody. This will help non-professional singers to produce more appealing karaoke performances. The quality of separated accompaniments was quantitatively evaluated with approximately one hour of polyphonic music, including material from a commercial karaoke DVD. Matti Ryynänen, Tuomas Virtanen, Jouni Paulus, Anssi Klapuri |
ICME | 2 |
| 2007 | Query by Example of Audio Signals using Euclidean Distance Between Gaussian Mixture ModelsabstractQuery by example of multimedia signals aims at automatic retrieval of media samples from a database, which are similar to a user-provided example. This paper proposes a method for query by example of audio signals. The method calculates a set of acoustic features from the signals and models their probability density functions (pdfs) using Gaussian mixture models. The method measures the similarity between two samples using the Euclidian distance between their pdfs. A novel method for calculating the closed form solution of the distance is proposed. Simulation experiments show that proposed method enables higher retrieval accuracy than the reference methods. Marko Leonard Helén, Tuomas Virtanen |
ICASSP (1) | 2 |
| 2007 | Monaural Sound Source Separation by Nonnegative Matrix Factorization With Temporal Continuity and Sparseness CriteriaabstractAn unsupervised learning algorithm for the separation of sound sources in one-channel music signals is presented. The algorithm is based on factorizing the magnitude spectrogram of an input signal into a sum of components, each of which has a fixed magnitude spectrum and a time-varying gain. Each sound source, in turn, is modeled as a sum of one or more components. The parameters of the components are estimated by minimizing the reconstruction error between the input spectrogram and the model, while restricting the component spectrograms to be nonnegative and favoring components whose gains are slowly varying and sparse. Temporal continuity is favored by using a cost term which is the sum of squared differences between the gains in adjacent frames, and sparseness is favored by penalizing nonzero gains. The proposed iterative estimation algorithm is initialized with random values, and the gains and the spectra are then alternatively updated using multiplicative update rules until the values converge. Simulation experiments were carried out using generated mixtures of pitched musical instrument samples and drum sounds. The performance of the proposed method was compared with independent subspace analysis and basic nonnegative matrix factorization, which are based on the same linear model. According to these simulations, the proposed method enables a better separation quality than the previous algorithms. Especially, the temporal continuity criterion improved the detection of pitched musical sounds. The sparseness criterion did not produce significant improvements Tuomas Virtanen |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Speech recognition using factorial hidden Markov models for separation in the feature spaceabstractThis paper proposes an algorithm for the recognition and separation of speech signals in non-stationary noise, such as another speaker. We present a method to combine hidden Markov models (HMMs) trained for the speech and noise into a factorial HMM to model the mixture signal. Robustness is obtained by separating the speech and noise signals in a feature domain, which discards unnecessary information. We use mel-cepstral coefficients (MFCCs) as features, and estimate the distribution of mixture MFCCs from the distributions of the target speech and noise. A decoding algorithm is proposed for finding the state transition paths and estimating gains for the speech and noise from a mixture signal. Simulations were carried out using speech material where two speakers were mixed at various levels, and even for high noise level (9 dB above the speech level), the method produced relatively good (60 % word recognition accuracy) results. Audio demonstrations are available at www.cs.tut.fi/˜tuomasv. Tuomas Virtanen |
INTERSPEECH | 1 |
| 2002 | Separation of harmonic sounds using linear models for the overtone seriesabstractA signal processing method is described, which separates harmonic sounds by applying linear models for the overtone series of sounds. Time-varying sinusoidal parameters are estimated in an iterative algorithm which is initialized using a multipitch estimator that finds the number of concurrent sounds and their frequency components. The iterative process then improves the estimates using the least-squares criterion. The harmonic stucture is retained by keeping the frequency ratio of overtones constant over time. Overlapping frequency components are resolved by using linear models for the overtone amplitudes. In practice, the models retain the spectral continuity of natural sounds. Simulation experiments were done using some basic structures for the linear models. These include polynomial, mel-cepstal and frequency-band model. Demonstration signals are available at http://www.cs.tut.fi/∼momasv/demopage.html. Tuomas Virtanen, Anssi Klapuri |
ICASSP | 1 |
| 2000 | Separation of harmonic sound sources using sinusoidal modelingabstractIn this paper, an approach for the separation of harmonic sounds is described. The overall system consists of three components. Sinusoidal modeling is first used to analyze the mixed signal and to obtain the frequencies and amplitudes of sinusoidal spectral components. Then a new method is proposed for the calculation of the perceptual distance between pairs of sinusoidal trajectories, according to the implications of psychoacoustic knowledge. A procedure for classifying the sinusoids into separate sound sources is presented. The system is not designed to separate sounds that have the same fundamental frequencies. However, a solution to detect single colliding sinusoids is given. Tuomas Virtanen, Anssi Klapuri |
ICASSP | 1 |