Sriram Srinivasan 0003

dblp:74/3399-3 · DBLP profile ↗
← Back
27ranked-venue papers
13as first author
9since 2021 · last 2026
0009-0004-5939-2368ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 23 · 9 first-author · 9 since 2021Artificial intelligence and machine learning · 14 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Echo Confidence: A Closed-Loop Metric for Residual Echo Monitoring and Dynamic Echo Control
George Gao, Chunmao Zhang, Puneet Rana, Sriram Srinivasan 0003
QoMEX4
2025 SMARTMOS: Modeling Subjective Audio Quality Evaluation for Real-Time Applications
Jose Antonio Jimenez Amador, Kaustubh Kalgaonkar, King-Wei Hor, Sriram Srinivasan 0003
INTERSPEECH5
2024 Interference Aware Training Target for DNN based joint Acoustic Echo Cancellation and Noise Suppression
Vahid Khanagha, Dimitris Koutsaidis, Kaustubh Kalgaonkar, Sriram Srinivasan 0003
INTERSPEECH4
2023 SCA: Streaming Cross-Attention Alignment For Echo Cancellation
abstract
End-to-End deep learning has shown promising results for speech enhancement tasks, such as noise suppression, dereverberation, and speech separation. However, most state-of-the-art methods for echo cancellation are either classical DSP-based or hybrid DSP-ML algorithms. Components such as the delay estimator and adaptive linear filter are based on traditional signal processing concepts, and deep learning algorithms typically only serve to replace the non-linear residual echo suppressor. This paper introduces an end-to-end echo cancellation network with a streaming cross-attention alignment (SCA). Our proposed method can handle unaligned inputs without requiring external alignment and generate high-quality speech without echoes. At the same time, the end-to-end algorithm simplifies the current echo cancellation pipeline for time-variant echo path cases. We test our proposed method on the ICASSP2022 and Inter-speech2021 Microsoft deep echo cancellation challenge evaluation dataset, where our method outperforms some of the other hybrid and end-to-end methods.
Yang Liu 0175, Yangyang Shi, Kaustubh Kalgaonkar, Sriram Srinivasan 0003
ICASSP5
2021 Interactive Speech and Noise Modeling for Speech Enhancement
abstract
Speech enhancement is challenging because of the diversity of background noise types. Most of the existing methods are focused on modelling the speech rather than the noise. In this paper, we propose a novel idea to model speech and noise simultaneously in a two-branch convolutional neural network, namely SN-Net. In SN-Net, the two branches predict speech and noise, respectively. Instead of information fusion only at the final output layer, interaction modules are introduced at several intermediate feature domains between the two branches to benefit each other. Such an interaction can leverage features learned from one branch to counteract the undesired part and restore the missing component of the other and thus enhance their discrimination capabilities. We also design a feature extraction module, namely residual-convolution-and-attention (RA), to capture the correlations along temporal and frequency dimensions for both the speech and the noises. Evaluations on public datasets show that the interaction module plays a key role in simultaneous modeling and the SN-Net outperforms the state-of-the-art by a large margin on various evaluation metrics. The proposed SN-Net also shows superior performance for speaker separation.
Xiulian Peng, Yuan Zhang 0013, Sriram Srinivasan 0003, Yan Lu 0001
AAAI4
2021 ICASSP 2021 Deep Noise Suppression Challenge
abstract
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH 2020 where we open-sourced training and test datasets for researchers to train their noise suppression models. We also open-sourced a subjective evaluation framework and used the tool to evaluate and select the final winners. Many researchers from academia and industry made significant contributions to push the field forward. We also learned that as a research community, we still have a long way to go in achieving excellent speech quality in challenging noisy real-time conditions. In this challenge, we expanded both our training and test datasets. Clean speech in the training set has increased by 200% with the addition of singing voice, emotion data, and non-English languages. The test set has increased by 100% with the addition of singing, emotional, non-English (tonal and non-tonal) languages, and, personalized DNS test clips. There are two tracks with focus on (i) real-time denoising, and (ii) real-time personalized DNS. We present the challenge results at the end.
Chandan K. A. Reddy, Harishchandra Dubey, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003
ICASSP8
2021 ICASSP 2021 Acoustic Echo Cancellation Challenge: Datasets, Testing Framework, and Results
abstract
The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real environments, as well as a synthetic dataset. We open source two large test sets, and we open source an online subjective test framework for researchers to quickly test their results. The winners of this challenge will be selected based on the average Mean Opinion Score (MOS) achieved across all different single talk and double talk scenarios.
Kusha Sridhar, Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Hannes Gamper, Sebastian Braun, Robert Aichner, Sriram Srinivasan 0003
ICASSP9
2021 INTERSPEECH 2021 Acoustic Echo Cancellation Challenge
Ross Cutler, Ando Saabas, Tanel Pärnamaa, Markus Loide, Sten Sootla, Marju Purin, Hannes Gamper, Sebastian Braun, Karsten Sørensen, Robert Aichner, Sriram Srinivasan 0003
Interspeech11
2021 INTERSPEECH 2021 Deep Noise Suppression Challenge
abstract
The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP 2020. We open-sourced training and test datasets for the wideband scenario. We also open-sourced a subjective evaluation framework based on ITU-T standard P.808, which was also used to evaluate participants of the challenge. Many researchers from academia and industry made significant contributions to push the field forward, yet even the best noise suppressor was far from achieving superior speech quality in challenging scenarios. In this version of the challenge organized at INTERSPEECH 2021, we are expanding both our training and test datasets to accommodate full band scenarios. The two tracks in this challenge will focus on real-time denoising for (i) wide band, and(ii) full band scenarios. We are also making available a reliable non-intrusive objective speech quality metric called DNSMOS for the participants to use during their development phase.
Chandan K. A. Reddy, Harishchandra Dubey, Kazuhito Koishida, Arun Asokan Nair, Vishak Gopal, Ross Cutler, Sebastian Braun, Hannes Gamper, Robert Aichner, Sriram Srinivasan 0003
Interspeech10
2020 DNN No-Reference PSTN Speech Quality Prediction
abstract
Classic public switched telephone networks (PSTN) are often a black box for VoIP network providers, as they have no access to performance indicators, such as delay or packet loss. Only the degraded output speech signal can be used to monitor the speech quality of these networks. However, the current state-of-the-art speech quality models are not reliable enough to be used for live monitoring. One of the reasons for this is that PSTN distortions can be unique depending on the provider and country, which makes it difficult to train a model that generalizes well for different PSTN networks. In this paper, we present a new open-source PSTN speech quality test set with over 1000 crowdsourced real phone calls. Our proposed no-reference model outperforms the full-reference POLQA and no-reference P.563 on the validation and test set. Further, we analyzed the influence of file cropping on the perceived speech quality and the influence of the number of ratings and training size on the model accuracy.
Gabriel Mittag, Ross Cutler, Yasaman Hosseinkashi, Michael Revow, Sriram Srinivasan 0003, Naglakshmi Chande, Robert Aichner
INTERSPEECH5
2020 The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
abstract
The INTERSPEECH 2020 Deep Noise Suppression (DNS) Challenge is intended to promote collaborative research in real-time single-channel Speech Enhancement aimed to maximize the subjective (perceptual) quality of the enhanced speech. A typical approach to evaluate the noise suppression methods is to use objective metrics on the test set obtained by splitting the original dataset. While the performance is good on the synthetic test set, often the model performance degrades significantly on real recordings. Also, most of the conventional objective metrics do not correlate well with subjective tests and lab subjective tests are not scalable for a large test set. In this challenge, we open-sourced a large clean speech and noise corpus for training the noise suppression models and a representative test set to real-world scenarios consisting of both synthetic and real recordings. We also open-sourced an online subjective test framework based on ITU-T P.808 for researchers to reliably test their developments. We evaluated the results using P.808 on a blind test set. The results and the key learnings from the challenge are discussed. The datasets and scripts can be found here for quick access https://github.com/microsoft/DNS-Challenge.
Chandan K. A. Reddy, Vishak Gopal, Ross Cutler, Ebrahim Beyrami, Roger Cheng, Harishchandra Dubey, Sergiy Matusevych, Robert Aichner, Ashkan Aazami, Sebastian Braun, Puneet Rana, Sriram Srinivasan 0003, Johannes Gehrke
INTERSPEECH12
2019 A Scalable Noisy Speech Dataset and Online Subjective Test Framework
abstract
Background noise is a major source of quality impairments in Voice over Internet Protocol (VoIP) and Public Switched Telephone Network (PSTN) calls.Recent work shows the efficacy of deep learning for noise suppression, but the datasets have been relatively small compared to those used in other domains (e.g., ImageNet) and the associated evaluations have been more focused.In order to better facilitate deep learning research in Speech Enhancement, we present a noisy speech dataset (MS-SNSD) that can scale to arbitrary sizes depending on the number of speakers, noise types, and Speech to Noise Ratio (SNR) levels desired.We show that increasing dataset sizes increases noise suppression performance as expected.In addition, we provide an open-source evaluation methodology to evaluate the results subjectively at scale using crowdsourcing, with a reference algorithm to normalize the results.To demonstrate the dataset and evaluation framework we apply it to several noise suppressors and compare the subjective Mean Opinion Score (MOS) with objective quality measures such as SNR, PESQ, POLQA, and VISQOL and show why MOS is still required.Our subjective MOS evaluation is the first large scale evaluation of Speech Enhancement algorithms that we are aware of.
Chandan K. A. Reddy, Ebrahim Beyrami, Jamie Pool, Ross Cutler, Sriram Srinivasan 0003, Johannes Gehrke
INTERSPEECH5
2014 Optimal rate allocation for speech enhancement using remote power-constrained wireless microphones
abstract
The problem of speech enhancement is considered in an interference environment, typical to applications like hands‐free voice communication and multi‐party conferencing. In the proposed system, a directional microphone placed at each interference is used to estimate the power spectral density (PSD) of the interference and the quantised PSD estimate is transmitted over a wireless link to an omnidirectional primary microphone that observes the corrupted source speech signal. At the primary microphone, the received observations are fused to obtain an estimate of the PSD of the desired speech signal. The problem of minimising transmitted power is considered subject to constraints on total transmission rate and maintaining the mean‐squared error in the estimated speech signal PSD below a prescribed limit, under narrowband and broadband signal models. The optimisation problem is solved by determining the optimum rate for encoding signal PSDs. The proposed strategy is analysed using sample speech and music signals.
Sriram Srinivasan 0003, Ashish Pandharipande
IET Signal Process.1
2013 Utility of auxiliary sensor data for speech enhancement
abstract
In recent years, data from various auxiliary acoustic and nonacoustic sensors have been used for enhancing noisy speech. These include bone-conduction microphones, surface electromyographic sensors, ultrasonic imaging of facial movements, etc. The signal from such sensors is correlated with the speech signal to varying degrees, and unlike microphone data, is typically not affected by acoustic background noise, making its use attractive for speech enhancement. In this paper, we discuss the measurement of the utility of such data from an information-theoretic perspective, and quantify the information that is shared between clean speech and the auxiliary signal, which is not present in the observed noisy speech signal. The measure is applied to simultaneously recorded air-and bone-conducted speech data.
Sriram Srinivasan 0003, Patrick Kechichian
ICASSP1
2012 Bit-rate reduction strategies for noise suppression with a remote wireless microphone
abstract
In single-channel non-stationary noise reduction it is paramount that a good noise reference is available in a timely manner to maintain a high quality speech signal. Using a remote wireless microphone placed close to a noise source, a good estimate of the noise power spectral density (PSD) can be acquired. This estimate, however, needs to be transmitted to the primary microphone for noise reduction. As wireless transmission is power intensive, it is desirable to reduce the bit-rate while maintaining good performance. In this paper, we propose techniques such as quantizing, frequency bin clubbing and intermittent PSD transmission to reduce the transmission bit-rate, and investigate their impact on performance.
Nemanja Cvijanovic, Ousman Sadiq, Sriram Srinivasan 0003
ICASSP3
2012 A Bayesian framework for robust speech enhancement under varying contexts
abstract
Single-microphone speech enhancement algorithms that employ trained codebooks of parametric representations of speech spectra have been shown to be successful in the suppression of non-stationary noise, e.g., in mobile phones. In this paper, we introduce the concept of a context-dependent codebook, and look at two aspects of context: dependency on the particular speaker using the mobile device, and on the acoustic condition during usage (e.g., hands-free mode in a reverberant room). Such context-dependent codebooks may be trained on-line. A new scheme is proposed to appropriately combine the estimates resulting from the context-dependent and context-independent codebooks under a Bayesian framework. Experimental results establish that the proposed approach performs better than the context-independent codebook in the case of a context match and better than the context-dependent codebook in the case of a context mismatch.
D. Hanumantha Rao Naidu, Sriram Srinivasan 0003
ICASSP2
2011 Using a remotewireless microphone for speech enhancement in non-stationary noise
abstract
As portable wireless audio-enabled devices become common, it is possible to form an ad-hoc network of such devices to enable high quality speech capture. In this paper, we consider the use of a remote wireless microphone placed close to a noise source, which transmits relevant information to the primary device, where it is used for noise reduction. Specific challenges introduced by such a scheme are addressed. It is seen that the proposed arrangement can result in a good amount of noise reduction in the presence of highly non stationary interferences such as music, where conventional single-microphone methods perform poorly. Improvements in segmental signal-noise-ratio of about 6-7 dB are observed when the method is applied to enhance noisy speech.
Sriram Srinivasan 0003
ICASSP1
2010 Ziv-Zakai bound for rate-constrained time-delay estimation in a wireless sensor network
Sriram Srinivasan 0003
FUSION1
2008 Spatial audio activity detection for hearing aids
abstract
We present a multi-microphone signal activity detection scheme for hearing aids to differentiate between the periods of activity of desired and interfering sources. The method is designed to provide robust performance in the presence of simultaneously active desired and interfering sources. We exploit knowledge from the hearing aid domain, and the directional processing present in modern hearing aids, to present a framework to design appropriate thresholds for the detection. Experiments confirm robust performance under practical reverberant conditions.
Sriram Srinivasan 0003, Kees Janse
ICASSP1
2007 Codebook-Based Bayesian Speech Enhancement for Nonstationary Environments
abstract
In this paper, we propose a Bayesian minimum mean squared error approach for the joint estimation of the short-term predictor parameters of speech and noise, from the noisy observation. We use trained codebooks of speech and noise linear predictive coefficients to model the a priori information required by the Bayesian scheme. In contrast to current Bayesian estimation approaches that consider the excitation variances as part of the a priori information, in the proposed method they are computed online for each short-time segment, based on the observation at hand. Consequently, the method performs well in nonstationary noise conditions. The resulting estimates of the speech and noise spectra can be used in a Wiener filter or any state-of-the-art speech enhancement system. We develop both memoryless (using information from the current frame alone) and memory-based (using information from the current and previous frames) estimators. Estimation of functions of the short-term predictor parameters is also addressed, in particular one that leads to the minimum mean squared error estimate of the clean speech signal. Experiments indicate that the scheme proposed in this paper performs significantly better than competing methods
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn
IEEE Trans. Speech Audio Process.1
2006 Multichannel parametric speech enhancement
abstract
We present a parametric model-based multichannel approach for speech enhancement. By employing an autoregressive model for the speech signal and using a trained codebook of speech linear predictive coefficients, minimum mean square error estimation of the speech signal is performed. By explicitly accounting for steering errors in the signal model, robust estimates are obtained. Experiments show that the proposed method results in significant performance gains.
Sriram Srinivasan 0003, Robert Aichner, W. Bastiaan Kleijn, Walter Kellermann
IEEE Signal Process. Lett.1
2006 Codebook driven short-term predictor parameter estimation for speech enhancement
abstract
In this paper, we present a new technique for the estimation of short-term linear predictive parameters of speech and noise from noisy data and their subsequent use in waveform enhancement schemes. The method exploits a priori information about speech and noise spectral shapes stored in trained codebooks, parameterized as linear predictive coefficients. The method also uses information about noise statistics estimated from the noisy observation. Maximum-likelihood estimates of the speech and noise short-term predictor parameters are obtained by searching for the combination of codebook entries that optimizes the likelihood. The estimation involves the computation of the excitation variances of the speech and noise auto-regressive models on a frame-by-frame basis, using the a priori information and the noisy observation. The high computational complexity resulting from a full search of the joint speech and noise codebooks is avoided through an iterative optimization procedure. We introduce a classified noise codebook scheme that uses different noise codebooks for different noise types. Experimental results show that the use of a priori information and the calculation of the instantaneous speech and noise excitation variances on a frame-by-frame basis result in good performance in both stationary and nonstationary noise conditions.
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn
IEEE Trans. Speech Audio Process.1
2005 Codebook-Based Bayesian Speech Enhancement
abstract
In this paper, we propose a Bayesian approach for the estimation of the short-term predictor parameters of speech and noise, from the noisy observation. The resulting estimates of the speech and noise spectra can be used in a Wiener filter or any state-of-the-art speech enhancement system. We utilize a-priori information about both speech and noise in the form of trained codebooks of linear predictive coefficients. In contrast to current Bayesian estimation approaches that consider the excitation variances as part of the a-priori information, in the proposed method they are computed analytically based on the observation at hand. Consequently, the method performs well in nonstationary noise conditions. Experimental results confirm the superior performance of the proposed method compared to existing Bayesian approaches, such as those based on hidden Markov models.
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn
ICASSP (1)1
2005 Denoising through source separation and minimum tracking
abstract
In this paper, we develop a multi-channel noise reduction algorithm based on blind source separation (BSS). In contrast to general BSS algorithms that attempt to recover all the signals, we explicitly estimate only the speech signal. By tracking the minimum of the spectral density of the microphone signals, noise-only segments are identified. The coefficients of the unmixing matrix that are necessary to separate the speech are identified from these segments through the optimization of an appropriate energy criterion. Since the proposed method explicitly estimates the speech signal from the noisy mixture, it does not suffer from the permutation problem that is typical to conventional BSS techniques. The method is applicable to both instantaneous and convolutive mixtures and achieves the separation in a single step, without the need for iterations. Experimental results show superior performance compared to a general BSS algorithm.
Sriram Srinivasan 0003, Mattias Nilsson 0002, W. Bastiaan Kleijn
INTERSPEECH1
2004 Estimation of short-term predictor parameters for coding and enhancement of noisy speech
abstract
We describe a technique for obtaining estimates of the short-term predictor parameters of speech under noisy conditions. We use a-priori information about speech in the form of a trained codebook of speech linear predictive coefficients. Our contribution is two-fold. First, we provide a framework where the standard vector quantization search to obtain the quantized linear predictive coefficients can be replaced by a maximum likelihood search, given the noisy observation, the speech codebook and an estimate of the noise. This results in an enhancement method that is integrated with parametric coders such as linear predictive analysis-by-synthesis coders. Second, we provide a scheme where the chosen vector is not restricted to be an element of the codebook. An interpolative search between the maximum likelihood estimate and its nearest neighbors in the codebook is used to improve the precision of the estimated parameters. Such a scheme is relevant when enhancement is considered separately from coding. Experimental results show improved performance for the proposed methods.
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn
ICASSP (1)1
2004 Speech enhancement using adaptive time-domain segmentation
abstract
In this paper, we investigate the benets of using an adaptive seg-mentation of the speech signal in speech enhancement. The adap-tive segmentation scheme divides the signal into the longest seg-ments within which stationarity is preserved, thus providing a good time-frequency resolution. The segmentation is performed with the help of an orthogonal library of local cosine bases using a compu-tationally efcient tree-structured best-basis search. We show that such an adaptive segmentation results in improved speech enhance-ment compared to a xed segmentation. The resulting enhanced speech is free from musical noise, without any additional smooth-ing. 1.
Sriram Srinivasan 0003, W. Bastiaan Kleijn
INTERSPEECH1
2003 Speech enhancement using a-priori information
Sriram Srinivasan 0003, Jonas Samuelsson, W. Bastiaan Kleijn
INTERSPEECH1