Robert Aichner

dblp:76/4449 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
10since 2021 · last 2023
0009-0000-6754-107XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 9 since 2021Artificial intelligence and machine learning · 8 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2023 PLCMOS - A Data-driven Non-intrusive Metric for The Evaluation of Packet Loss Concealment Algorithms
Lorenz Diener, Marju Purin, Sten Sootla, Ando Saabas, Robert Aichner, Ross Cutler
INTERSPEECH5
2022 ICASSP 2022 Acoustic Echo Cancellation Challenge
abstract
The ICASSP 2022 Acoustic Echo Cancellation Challenge is intended to stimulate research in acoustic echo cancellation (AEC), which is an important area of speech enhancement and still a top issue in audio communication. This is the third AEC challenge and it is enhanced by including mobile scenarios, adding speech recognition word accuracy rate as a metric, and making the audio 48 kHz. We open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 10,000 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source an online subjective test framework and provide an online objective metric service for researchers to quickly test their results. The winners of this challenge were selected based on the average Mean Opinion Score (MOS) achieved across all scenarios and the word accuracy rate.
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner
ICASSP8
2022 Icassp 2022 Deep Noise Suppression Challenge
abstract
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. This is the 4th DNS challenge, with the previous editions held at INTERSPEECH 2020 [1], ICASSP 2021 [2], and INTERSPEECH 2021 [3]. We open-source datasets and test sets for researchers to train their deep noise suppression models, as well as a subjective evaluation framework based on ITU-T P.835 to rate and rank-order the challenge entries. We provide access to DNS-MOS P.835 and word accuracy (WAcc) APIs to challenge participants to help with iterative model improvements. In this challenge, we introduced the following changes: (i) Included mobile device scenarios in the blind test set; (ii) Included a personalized noise suppression track with baseline; (iii) Added WAcc as an objective metric; (iv) Included DNSMOS P.835; (v) Made the training datasets and test sets fullband (48 kHz). We use an average of WAcc and subjective scores P.835 SIG, BAK, and OVRL to get the final score for ranking the DNS models. We believe that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-world scenarios.
Harishchandra Dubey, Vishak Gopal, Ross Cutler, Ashkan Aazami, Sergiy Matusevych, Sebastian Braun, Sefik Emre Eskimez, Manthan Thakker, Takuya Yoshioka, Hannes Gamper, Robert Aichner
ICASSP11
2022 INTERSPEECH 2022 Audio Deep Packet Loss Concealment Challenge
abstract
Audio Packet Loss Concealment (PLC) is the hiding of gaps in audio streams caused by data transmission failures in packet switched networks.This is a common problem, and of increasing importance as end-to-end VoIP telephony and teleconference systems become the default and ever more widely used form of communication in business as well as in personal usage.This paper presents the INTERSPEECH 2022 Audio Deep Packet Loss Concealment challenge.We first give an overview of the PLC problem, and introduce some classical approaches to PLC as well as recent work.We then present the open source dataset released as part of this challenge as well as the evaluation methods and metrics used to determine the winner.We also briefly introduce PLCMOS, a novel data-driven metric that can be used to quickly evaluate the performance PLC systems.Finally, we present the results of the INTERSPEECH 2022 Audio Deep PLC Challenge, and provide a summary of important takeaways.
Lorenz Diener, Sten Sootla, Solomiya Branets, Ando Saabas, Robert Aichner, Ross Cutler
INTERSPEECH5
2022 MusicNet: Compact Convolutional Neural Network for Real-time Background Music Detection
abstract
With the recent growth of remote work, online meetings often encounter challenging audio contexts such as background noise, music, and echo.Accurate real-time detection of music events can help to improve the user experience.In this paper, we present MusicNet, a compact neural model for detecting background music in the real-time communications pipeline.In video meetings, music frequently co-occurs with speech and background noises, making the accurate classification quite challenging.We propose a compact convolutional neural network core preceded by an in-model featurization layer.MusicNet takes 9 seconds of raw audio as input and does not require any model-specific featurization in the product stack.We train our model on the balanced subset of the Audio Set [1] data and validate it on 1000 crowd-sourced real test clips.Finally, we compare MusicNet performance with 20 state-of-the-art models.MusicNet has a true positive rate (TPR) of 81.3% at a 0.1% false positive rate (FPR), which is significantly better than state-of-the-art models included in our study.MusicNet is also 10x smaller and has 4x faster inference than the best performing models we benchmarked.
Chandan K. A. Reddy, Vishak Gopal, Harishchandra Dubey, Ross Cutler, Sergiy Matusevych, Robert Aichner
INTERSPEECH6
2021 ICASSP 2021 Deep Noise Suppression Challenge
abstract
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to train their noise suppression models. We also open-sourced a subjective evaluation framework and used the tool to evaluate and select the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we expanded both our training and test datasets. Clean speech in the training set has increased by 200% with the addition of singing voice, emotion data, and non-English languages. The test set has increased by 100% with the addition of singing, emotional, non-English (tonal and non-tonal) languages, and, personalized DNS test clips. There are two tracks with focus on (i) real-time denoising, and (ii) real-time personalized DNS. We present the challenge results at the end.
Chandan K. A. Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003
ICASSP7
2021 ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results
abstract
The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source two large test sets, and we open source an online subjective test framework for researchers to quickly test their results. The winners of this challenge will be selected based on the average Mean Opinion Score (MOS) achieved across all different single talk and double talk scenarios.
Kusha Sridhar, Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Hannes Gamper, Sebastian Braun, Robert Aichner, Sriram Srinivasan 0003
ICASSP8
2021 INTERSPEECH 2021 Acoustic Echo Cancellation Challenge
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Sten Sootla, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner, Sriram Srinivasan 0003
Interspeech10
2021 INTERSPEECH 2021 Deep Noise Suppression Challenge
abstract
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase.
Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003
Interspeech9
2021 Meeting Effectiveness and Inclusiveness in Remote Collaboration
abstract
A primary goal of remote collaboration tools is to provide effective and inclusive meetings for all participants. To study meeting effectiveness and meeting inclusiveness, we first conducted a large-scale email survey (N=4,425; after filtering N=3,290) at a large technology company (pre-COVID-19); using this data we derived a multivariate model of meeting effectiveness and show how it correlates with meeting inclusiveness, participation, and feeling comfortable to contribute. We believe this is the first such model of meeting effectiveness and inclusiveness. The large size of the data provided the opportunity to analyze correlations that are specific to sub-populations such as the impact of video. The model shows the following factors are correlated with inclusiveness, effectiveness, participation, and feeling comfortable to contribute in meetings: sending a pre-meeting communication, sending a post-meeting summary, including a meeting agenda, attendee location, remote-only meeting, audio/video quality and reliability, video usage, and meeting size. The model and survey results give a quantitative understanding of how and where to improve meeting effectiveness and inclusiveness and what the potential returns are. Motivated by the email survey results, we implemented a post-meeting survey into a leading computer-mediated communication (CMC) system to directly measure meeting effectiveness and inclusiveness (during COVID-19). Using initial results based on internal flighting we created a similar model of effectiveness and inclusiveness, with many of the same findings as the email survey. This shows a method of measuring and understanding these metrics which are both practical and useful in a commercial CMC system. By improving meeting effectiveness, companies can save significant time and money. Improving meeting inclusiveness is hypothesized to improve meeting effectiveness, but also improves the working environment and employee retention at organizations.
Ross Cutler, Yasaman Hosseinkashi, Jamie Pool, Senja Filipi, Robert Aichner, Yuan Tu, Johannes Gehrke
Proc. ACM Hum. Comput. Interact.5
2020 DNN No-Reference PSTN Speech Quality Prediction
abstract
Classic public switched telephone networks (PSTN) are often a black box for VoIP network providers, as they have no access to performance indicators, such as delay or packet loss. Only the degraded output speech signal can be used to monitor the speech quality of these networks. However, the current state-of-the-art speech quality models are not reliable enough to be used for live monitoring. One of the reasons for this is that PSTN distortions can be unique depending on the provider and country, which makes it difficult to train a model that generalizes well for different PSTN networks. In this paper, we present a new open-source PSTN speech quality test set with over 1000 crowdsourced real phone calls. Our proposed no-reference model outperforms the full-reference POLQA and no-reference P.563 on the validation and test set. Further, we analyzed the influence of file cropping on the perceived speech quality and the influence of the number of ratings and training size on the model accuracy.
Gabriel Mittag, Ross Cutler, Yasaman Hosseinkashi, Michael Revow, Sriram Srinivasan 0003, Naglakshmi Chande, Robert Aichner
INTERSPEECH7
2020 The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
abstract
The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical approach to evaluate the noise suppression methods is to use objective metrics on the test set obtained by splitting the original dataset. While the performance is good on the synthetic test set, often the model performance degrades significantly on real recordings. Also, most of the conventional objective metrics do not correlate well with subjective tests and lab subjective tests are not scalable for a large test set. In this challenge, we open-sourced a large clean speech and noise corpus for training the noise suppression models and a representative test set to real-world scenarios consisting of both synthetic and real recordings. We also open-sourced an online subjective test framework based on ITU-T P.808 for researchers to reliably test their developments. We evaluated the results using P.808 on a blind test set. The results and the key learnings from the challenge are discussed. The datasets and scripts can be found here for quick access https://github.com/microsoft/DNS-Challenge.
Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan 0003, Johannes Gehrke
INTERSPEECH8
2007 Multi-Channel Source Separation Preserving Spatial Information
abstract
In this paper we propose two novel methods for preserving the spatial information in source separation algorithms. Our approach is applicable to any source separation algorithm and is based on an additional supervised adaptive filtering with the reference signals generated by the source separation system. If a special constrained optimization scheme is applied to derive the source separation algorithm then the novel approach can be simplified. The quality of the spatial representation and the separation performance of both methods and two state-of-the-art approaches from the literature have been evaluated by a MUSHRA listening test according to the relevant ITU recommendation showing that the novel methods clearly outperform the state-of-the-art approaches.
Robert Aichner, Herbert Buchner, Meray Zourub, Walter Kellermann
ICASSP (1)1
2006 Post-Processing for Convolutive Blind Source Separation
abstract
Convolutive blind source separation (BSS) aims at separating point sources from mixtures picked up by several sensors. In real-world environments moving speakers, background noise and long reverberation are encountered which often degrade the performance of BSS algorithms. In such cases, the application of a post-filter can improve the output signal quality by suppression of residual cross-talk and of background noise. In this paper we propose a novel technique to estimate the necessary power spectral densities of the cross-talk components and present a robust system which allows to further suppress both, the remaining interference from point sources and the background noise. Experimental results show the benefit of this post-processing method in realistic environments
Robert Aichner, Meray Zourub, Herbert Buchner, Walter Kellermann
ICASSP (5)1
2006 Separating Convolutive Mixtures with Trinicon
abstract
Blind source separation (BSS) algorithms are often categorized as either narrowband or broadband algorithms depending on whether their respective cost functions aim at individual DFT bins or the entire broadband signal. In this contribution, we present comparable general natural gradient-based formulations of both concepts based on the TRINICON framework. As a distinctive feature, narrowband algorithms imply an internal permutation and scaling problem. We show that the common DOA estimation-based methods for aligning the permutations effectively rely on geometric a-priori knowledge, and we explain why they need to be complemented by additional repair mechanisms for robust BSS. The latter can already be viewed as approximations of the generic TRINICON broadband algorithm. As a conclusion, we propose to always use a generic broadband algorithm as a starting point for the design of new BSS algorithms
Walter Kellermann, Herbert Buchner, Robert Aichner
ICASSP (5)3
2006 A real-time blind source separation scheme and its application to reverberant and noisy acoustic environments
Robert Aichner, Herbert Buchner, Walter Kellermann
Signal Process.1
2006 Multichannel parametric speech enhancement
abstract
We present a parametric model-based multichannel approach for speech enhancement. By employing an autoregressive model for the speech signal and using a trained codebook of speech linear predictive coefficients, minimum mean square error estimation of the speech signal is performed. By explicitly accounting for steering errors in the signal model, robust estimates are obtained. Experiments show that the proposed method results in significant performance gains.
Sriram Srinivasan 0003, Robert Aichner, W. Bastiaan Kleijn, Walter Kellermann
IEEE Signal Process. Lett.2
2005 On the causality problem in time-domain blind source separation and deconvolution algorithms
abstract
Using a recently presented generic framework for multichannel blind signal processing for convolutive mixtures, we investigate the problem of incorporating acausal delays which are necessary with certain geometric constellations. Starting from a generic update equation which is applicable to blind source separation (BSS), multichannel blind deconvolution (MCBD), and multichannel blind partial deconvolution (MCBPD) for dereverberation of speech signals, two formulations of the natural gradient are derived. It is shown that one expression is applicable to mere causal filters whereas the other also allows an implementation of noncausal filters. Moreover, proper initialization methods for both cases are given. For the implementation of these algorithms, cross-relation estimation techniques, known from linear prediction, are discussed. Based on these results, relationships between traditional MCBD algorithms can be established. Experimental results of different acoustic scenarios show the applicability of the presented algorithms.
Robert Aichner, Herbert Buchner, Walter Kellermann
ICASSP (5)1
2005 Simultaneous localization of multiple sound sources using blind adaptive MIMO filtering
abstract
Blind adaptive filtering for time delay of arrival (TDOA) estimation is a very powerful method for acoustic source localization in reverberant environments with broadband signals like speech. Based on a recently presented generic framework for blind signal processing for convolutive mixtures, called TRINICON, we present a TDOA estimation method for simultaneous multidimensional localization of multiple sources. Moreover, an interesting link to the known single-input multiple-output (SIMO)-based adaptive eigenvalue decomposition (AED) method is shown. We evaluate the novel multiple-input multiple-output (MIMO)-based approach and compare it with the known SIMO-based method in a reverberant acoustic environment using reference data of the positions obtained from infrared sensors. The results show that the new approach is very robust against reverberation and background noise.
Herbert Buchner, Robert Aichner, Jochen Stenglein, Heinz Teutsch, Walter Kellermann
ICASSP (3)2
2005 A generalization of blind source separation algorithms for convolutive mixtures based on second-order statistics
abstract
We present a general broadband approach to blind source separation (BSS) for convolutive mixtures based on second-order statistics. This avoids several known limitations of the conventional narrowband approximation, such as the internal permutation problem. In contrast to traditional narrowband approaches, the new framework simultaneously exploits the nonwhiteness property and nonstationarity property of the source signals. Using a novel matrix formulation, we rigorously derive the corresponding time-domain and frequency-domain broadband algorithms by generalizing a known cost-function which inherently allows joint optimization for several time-lags of the correlations. Based on the broadband approach time-domain, constraints are obtained which provide a deeper understanding of the internal permutation problem in traditional narrowband frequency-domain BSS. For both the time-domain and the frequency-domain versions, we discuss links to well-known, and also, to novel algorithms that constitute special cases. Moreover, using the so-called generalized coherence, links between the time-domain and the frequency-domain algorithms can be established, showing that our cost function leads to an update equation with an inherent normalization ensuring a robust adaptation behavior. The concept is applicable to offline, online, and block-online algorithms by introducing a general weighting function allowing for tracking of time-varying real acoustic environments.
Herbert Buchner, Robert Aichner, Walter Kellermann
IEEE Trans. Speech Audio Process.2
2004 TRINICON: a versatile framework for multichannel blind signal processing
abstract
In this paper we present a framework for multichannel blind signal processing for convolutive mixtures, such as blind source separation (BSS) and multichannel blind deconvolution (MCBD). It is based on the use of multivariate pdf and a compact matrix notation which considerably simplifies the representation and handling of the algorithms. By introducing these techniques into an information theoretic cost function, we can exploit the three fundamental signal properties nonwhiteness, nongaussianity, and nonstationarity. This results in a versatile tool that we call TRINICON (Triple-N ICA for convolutive mixtures). Both, links to popular algorithms and several novel algorithms follow from the general approach. In particular, we introduce a new concept of multichannel blind partial deconvolution (MCBPD) for speech which prevents a complete whitening of the output signals, i.e., the vocal tract is excluded from the equalization. This is especially interesting for automatic speech recognition applications. Moreover, we show results for BSS using multivariate spherically invariant random processes (SIRP) to efficiently model speech, and show how the approach carries over to MCBPD. These concepts are also suitable for an efficient implementation in the frequency domain by using a rigorous broadband derivation avoiding the internal permutation problem and circularity effects.
Herbert Buchner, Robert Aichner, Walter Kellermann
ICASSP (3)2
2003 Subband based blind source separation for convolutive mixtures of speech
abstract
Subband processing is applied to blind source separation (BSS) for convolutive mixtures of speech. This is motivated by the drawback of frequency-domain BSS, i.e., when a long frame with a fixed frame-shift is used to cover reverberation, the number of samples in each frequency decreases and the separation performance is degraded. In our proposed subband BSS, (1) by using a moderate number of subbands, a sufficient number of samples can be held in each subband, mid (2) by using FIR filters in each subband, we can handle long reverberation. Subband BSS achieves better performance than frequency-domain BSS. Moreover, we propose efficient separation procedures that take into consideration the frequency characteristics of room reverberation and speech signals. We achieve this (3) by using longer unmixing filters in low frequency bands, and (4) by adopting overlap-blockshift in BSS's batch adaptation in low frequency bands. Consequently, frequency-dependent subband processing is successfully realized in the proposed subband BSS.
Shoko Araki, Shoji Makino, Robert Aichner, Tsuyoki Nishikawa, Hiroshi Saruwatari
ICASSP (5)3