VLDB 2026 Research / reviewers in the wild / expert
Rainer Martin 0001
dblp:18/4464
· DBLP profile ↗
99ranked-venue papers
14as first author
14since 2021 · last 2024
0000-0002-9587-4215ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 76 · 12 first-author · 12 since 2021Artificial intelligence and machine learning · 41 · 6 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Sketching-Based Acoustic Scene Change Detection in Low-Power Embedded DevicesabstractAcoustic scene classification (ASC) in embedded devices (EDs) is a challenging task due to a large variety of acoustic environments occurring in daily life. In particular, ASC is error-prone in scenes which are not well represented in its training data. In ultra low-power EDs such as hearing aids, the performance of an embedded classifier is additionally constrained by limited computational resources. This calls for the development of classifiers that are tailored to the specific environment of individual users. Therefore, we propose an online scene change detection (SCD) framework that aggregates statistical information of extracted low-dimensional audio features in the ED and detects change-points such that data of previously unseen scenes can be aggregated in the ED and further be evaluated and used for re-training in a cloud-based system. We compare an approach based on second-order statistics with a method based on sketching and show that both achieve significantly better overall SCD performance with respect to a low-complexity baseline. Furthermore, the sketching-based approach provides a better detection-error trade-off than the method relying on second-order statistics only. Timm Koppelmann, Rainer Martin 0001 |
MMSP | 2 |
| 2024 | Direction and Distance Estimation of Whistle Events on a NAO Robot
Diana Kleingarn, Dominik Brämer, Rainer Martin 0001 |
RoboCup | 3 |
| 2024 | Conditional Label Smoothing For LLM-Based Data Augmentation in Medical Text ClassificationabstractIn this paper, we propose and evaluate a data augmentation technique for improving the classification of text data. Using a large language model (LLM), samples of real-world data and a tailored prompt, we generate additional text samples and use these in combination with real-world data to train a deep neural network (DNN) for the classification of medical sentences. In order to compensate for variations in the augmented data, we apply label smoothing (LS) to render the DNN more robust against mislabeled sentences. Moreover, we introduce a conditional label smoothing (CLS) approach to exploit the classification statistics. CLS re-parameterizes label noise, effectively reducing inter-class confusion. We evaluate our system on privacy-sensitive medical text data where we first show the general benefits of using LS in data augmentation and secondly compare CLS to LS. While LS already leads to an improved F1-score and a better false negative rate, CLS slightly outperforms ordinary LS. Luca Becker, Philip Pracht, Peter Sertdal, Jil Uboreck, Alexander Bendel, Rainer Martin 0001 |
SLT | 6 |
| 2023 | Head movements in two- and four-person interactive conversational tasks in noisy and moderately reverberant conditions
Alan Archer-Boyd, Rainer Martin 0001 |
INTERSPEECH | 2 |
| 2023 | Personalized Acoustic Scene Classification in Ultra-low Power Embedded Devices Using Privacy-preserving Data Augmentation
Timm Koppelmann, Semih Agcaer, Rainer Martin 0001 |
INTERSPEECH | 3 |
| 2023 | Short-term Extrapolation of Speech Signals Using Recursive Neural Networks in the STFT Domain
Maurice Oberhag, Daniel Neudek, Rainer Martin 0001, Tobias Rosenkranz, Henning Puder |
INTERSPEECH | 3 |
| 2022 | On Spectral and Temporal Sparsification of Speech Signals for the Improvement of Speech Perception in CI ListenersabstractThe perception of complex signals such as music or speech in noise is a difficult task for most cochlear implant (CI) users. Furthermore, there is a wide variability in speech recognition so that some also face difficulties in everyday situations with only little or no noise. In this study, two methods inspired by music simplification approaches were developed and evaluated through instrumental measures and in listening tests with adult CI listeners. Signals were processed based on a separation into transient and harmonic parts. After sparsification of the harmonic spectrum by principal component analysis (PCA) or by an individualized spectral peak picking approach, the transient and the sparsened harmonic parts were remixed. Significant improvements in speech recognition could be observed when the PCA-based method was applied. This might be caused by noise reduction effects and modulation amplifications as shown by correlation analysis. Benjamin Lentz, Rainer Martin 0001, Kirsten Oberländer, Christiane Völter |
ICASSP | 2 |
| 2022 | Clustering-based Wake Word Detection in Privacy-aware Acoustic Sensor NetworksabstractThis work investigates privacy-aware collaborative wake word detection (WWD) in acoustic sensor networks. To meet state-of-the-art privacy constraints, the proposed WWD scheme is based on privacy-aware unsupervised clustered federated learning that groups microphone nodes w.r.t. active sound sources and on a privacy-preserving high-level feature representation. Using the partition of microphone nodes into clusters, we apply intra- and inter-cluster feature enhancement strategies directly in the privacy-preserving feature domain and thus circumvent the need for communicating privacy-sensitive information between nodes. The approach is demonstrated for an acoustic sensor network deployed in a smart-home environment. We show that the proposed collaborative WWD system clearly outperforms independent decisions of individual microphone nodes. Index Terms: privacy, wake word detection, clustering, federated learning, unsupervised clustered federated learning Timm Koppelmann, Luca Becker, Alexandru Nelus, Rene Glitza, Lea Schönherr, Rainer Martin 0001 |
INTERSPEECH | 6 |
| 2022 | Spectral sparsification of speech signals and its interaction with top-down mechanisms in adult cochlear implant users
Benjamin Lentz, Christiane Völter, Rainer Martin 0001 |
Speech Commun. | 3 |
| 2022 | First-Order Recursive Smoothing of Short-Time Power Spectra in the Presence of InterferenceabstractIn this letter we derive an optimal time-varying smoothing factor for smoothing non-stationary power spectra of a target signal when a noisy observation of this signal is given. The proposed approach is based on a complex Gaussian signal model in the short-time discrete Fourier domain and on mean squared error optimization. We find that the optimal smoothing factor depends on the signal-to-noise ratio as well as on the deviation between the smoothed estimate and the target signal power spectra. We investigate the properties of the proposed adaptive smoothing system and demonstrate its utility in a basic validation experiment. Jalal Taghia, Daniel Neudek, Tobias Rosenkranz, Henning Puder, Rainer Martin 0001 |
IEEE Signal Process. Lett. | 5 |
| 2021 | A DNN Autoencoder for Automotive Radar Interference MitigationabstractIn this paper, a novel interference mitigation approach using an autoencoder in combination with a traditional interference detection filter is introduced. It is shown that by employing the gated convolution, the encoder has the ability to learn the signal pattern from the remaining interference-free signal. The decoder can recover the interference-contaminated signal segments from the bottleneck representation as computed by the encoder. Experimental results show that the proposed method can provide a remarkable improvement in signal-to-interference-plus-noise ratio (SINR) and preserves its robustness on real radar measurements in severely disturbed scenarios that are more complex than the training dataset. Shengyi Chen, Jalal Taghia, Tai Fei, Uwe Kühnau, Nils Pohl, Rainer Martin 0001 |
ICASSP | 6 |
| 2021 | Estimation of Microphone Clusters in Acoustic Sensor Networks Using Unsupervised Federated LearningabstractIn this paper we present a privacy-aware method for estimating source-dominated microphone clusters in the context of acoustic sensor networks (ASNs). The approach is based on clustered federated learning which we adapt to unsupervised scenarios by employing a light-weight autoencoder model. The model is further optimized for training on very scarce data. In order to best harness the benefits of clustered microphone nodes in ASN applications, a method for the computation of cluster membership values is introduced. We validate the performance of the proposed approach using clustering-based measures and a network-wide classification task. Alexandru Nelus, Rene Glitza, Rainer Martin 0001 |
ICASSP | 3 |
| 2021 | Privacy-Preserving Feature Extraction for Cloud-Based Wake Word VerificationabstractWake word detection and verification systems often involve a local, on-device wake word detector and a cloud-based verification node. In such systems, the audio representation sent to the cloud-based server may exhibit sensitive information that might be intercepted by an eavesdropper. To improve privacy of cloud-based wake word verification (WWV) systems, we propose to use a privacy-preserving feature representation that minimizes the automatic speech recognition (ASR) capability of a potential attacker. The proposed approach employs an adversarial training schedule that aims to minimize an attacker’s word error rate (WER) while maintaining a high WWV performance. To this end, we apply an adaptive weighting factor in the combined loss function to control the balance between minimizing the WWV loss and maximizing the ASR loss. We show that the proposed training method significantly reduces possible privacy risks while maintaining a strong WWV performance. Timm Koppelmann, Alexandru Nelus, Lea Schönherr, Dorothea Kolossa, Rainer Martin 0001 |
Interspeech | 5 |
| 2021 | Privacy-Preserving Audio Classification Using Variational Information Feature ExtractionabstractIn this paper we investigate and tackle the privacy risks of deep-neural-network-based feature extraction for sound classification in acoustic sensor networks. To this end, we analyze a single-label domestic activity monitoring and a multi-label urban sound tagging scenario. We show that in both cases, the feature representations designed for sound classification also carry a significant amount of speaker-dependent data, thus posing serious privacy risks for speaker recognition attacks based on feature interception. We then propose to mitigate the aforementioned privacy risks by introducing a variational information feature extraction scheme that allows sound classification while, concurrently, minimizing the feature representation's level of information and hence, inhibiting speaker recognition attempts. We control and analyze the balance between the performance of the trusted and attacker tasks via the resulting model's composite loss function, its budget scaling factor, and latent space size. It is empirically demonstrated that the proposed privacy-preserving feature representation generalizes well to both single-label and multi-label scenarios with vast as well as reduced training-dataset resources. Furthermore, it exhibits robustness against x-vector-based, state-of-the-art speaker recognition attacks. Alexandru Nelus, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Audio Feature Extraction for Vehicle Engine Noise ClassificationabstractIn this paper we propose a new scheme for vehicle engine noise classification as a more privacy-preserving alternative to classifying vehicles based on video recordings. We establish two scenarios: diesel vs. petrol and heavy goods vehicle vs. personal car classification. Our approach includes a novel modulation-spectrum-based feature representation that is used in conjunction with a siamese neural network classifier. Additionally, a database containing recordings from diverse urban acoustic scenarios is provided. The obtained results show the advantage of the proposed approach compared to conventional feature representations and classifiers. This is achieved by de-correlating background noise from target noise and by quantifying the degree of variation of noise characteristics. Luca Becker, Alexandru Nelus, Johannes Gauer, Lars Rudolph, Rainer Martin 0001 |
ICASSP | 5 |
| 2020 | Harmonic/Percussive Sound Separation and Spectral Complexity Reduction of Music Signals for Cochlear Implant ListenersabstractCochlear implant (CI) users suffer from limitations in music perception and thus prefer music which has a clear rhythm/beat and is played with only a few instruments. Therefore, existing music pre-processing methods aim to enhance music signals for CI users by either emphasizing preferred voices or reducing the spectral complexity of the signals. In this work, a music pre-processing scheme is described which combines these approaches and is applicable to a wider variety of music genres. The proposed method is evaluated and compared to other recently developed methods using instrumental measures and a listening test with vocoded pop/rock music excerpts and normal hearing listeners. Unprocessed popular music pieces as well as different processed versions were rated comparatively in terms of distinctness of drums, distinctness of melody, and the overall impression. The listening test showed significantly better ratings for the proposed method compared to unprocessed music and most of the other processing schemes. As the instrumental measures also indicate improvements, the proposed combined strategy is a promising candidate for music enhancement for CI listeners. Benjamin Lentz, Anil M. Nagathil, Johannes Gauer, Rainer Martin 0001 |
ICASSP | 4 |
| 2020 | Binaural Direct-to-Reverberant Energy Ratio and Speaker Distance EstimationabstractThis article addresses the problem of distance estimation using binaural hearing aid microphones in reverberant rooms. Among several distance indicators, the direct-to-reverberant energy ratio (DRR) has been shown to be more effective than other features. Therefore, we present two novel approaches to estimate the DRR of binaural signals. The first method is based on the interaural magnitude-squared coherence whereas the second approach uses stochastic maximum likelihood beamforming to estimate the power of the direct and reverberant components. The proposed DRR estimation algorithms are integrated into a distance estimation technique. When based solely on DRR, the distance estimation algorithm requires calibration where naturally the critical distance is a good calibration point. We thus propose two approaches for the calibration of the distance estimation algorithm: Informed calibration using the critical distance of the reverberant room and blind calibration using the listener's own voice. Results across various acoustical environments show the benefit of the proposed algorithms for the estimation of sound source distances up to 3 m with an estimation error of about 35 cm using informed calibration and about 1 m using the fully blind calibration strategy. Mehdi Zohourian, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Privacy-aware Feature Extraction for Gender Discrimination versus Speaker IdentificationabstractThis paper introduces a deep neural network based feature extraction scheme that aims to improve the trade-off between utility and privacy in speaker classification tasks. In the proposed scenario we develop a feature representation that helps to maximize the performance of a gender classifier while minimizing additional speaker identity information. Our approach is to use variational information feature extraction that allows for gender discrimination (utility) but minimizes the information level of the features, thus discouraging speaker identification adversarial attacks (privacy). We analyze the model's loss function and the budget scaling factor used to control the balance of utility vs. privacy. It is experimentally shown that the proposed method reduces privacy risks without significantly deprecating utility and that it also generalizes well to new speaker contexts. Alexandru Nelus, Rainer Martin 0001 |
ICASSP | 2 |
| 2019 | Direct-to-reverberant Energy Ratio Estimation Based on Interaural Coherence and a Joint ITD/ILD ModelabstractThis paper proposes a novel algorithm to estimate the direct-to-reverberant energy ratio (DRR) using hearing aid microphones. The algorithm is based on the interaural magnitude-squared coherence of signals and is able to take both phase and level differences of microphones signals in the binaural configuration into account. We employ a spherical head model to approximate binaural cues. The proposed algorithm uses the common assumption of an ideally diffuse reverberation sound field. We test our approach on signals based on simulated and on measured binaural room impulse responses. Results show improved performance of the proposed algorithm as compared to other coherence-based DRR estimation methods. Mehdi Zohourian, Rainer Martin 0001 |
ICASSP | 2 |
| 2019 | Privacy-Preserving Variational Information Feature Extraction for Domestic Activity Monitoring versus Speaker IdentificationabstractIn this paper we highlight the privacy risks entailed in deep neural network feature extraction for domestic activity monitoring. We employ the baseline system proposed in the Task 5 of the DCASE 2018 challenge and simulate a feature interception attack by an eavesdropper who wants to perform speaker identification. We then propose to reduce the aforementioned privacy risks by introducing a variational information feature extraction scheme that allows for good activity monitoring performance while at the same time minimizing the information of the feature representation, thus restricting speaker identification attempts. We analyze the resulting model’s composite loss function and the budget scaling factor used to control the balance between the performance of the trusted and attacker tasks. It is empirically demonstrated that the proposed method reduces speaker identification privacy risks without significantly deprecating the performance of domestic activity monitoring tasks. Alexandru Nelus, Janek Ebbers, Reinhold Häb-Umbach, Rainer Martin 0001 |
INTERSPEECH | 4 |
| 2019 | Privacy-Preserving Siamese Feature Extraction for Gender Recognition versus Speaker Identification
Alexandru Nelus, Silas Rech, Timm Koppelmann, Henrik Biermann, Rainer Martin 0001 |
INTERSPEECH | 5 |
| 2018 | Binaural Spectral Complexity Reduction of Music Signals for Cochlear Implant ListenersabstractAn emphasis on the leading voice or melody is known to facilitate music perception in cochlear implant (CI) listeners while a competing accompaniment is perceived as disturbing. In this paper we present the extension of a monaural music complexity reduction scheme for CI users towards a binaural application. The scheme aims at an attenuation of the accompaniment in music signals and relies on a reduced-rank approximation by means of principal component analysis (PCA). In the proposed binaural system the PCA is only performed for the melody dominated ear and its eigenvectors are used for a reduced-rank representation of both ear signals. We use SIR and SAR measures for evaluation and show that with binaural processing a further attenuation of the accompaniment can be achieved in comparison to separate bilateral processing of both ear signals. At the same time neither additional artifacts are introduced to the reconstructed melody signals nor the binaural cues accessible to CI users are considerably harmed. Johannes Gauer, Anil M. Nagathil, Rainer Martin 0001 |
ICASSP | 3 |
| 2018 | GSC-Based Binaural Speaker Separation Preserving Spatial CuesabstractIn this paper we investigate two methods for the preservation of spatial cues in binaural speaker separation. We develop these methods as extensions of our previously proposed model-based generalized sidelobe canceller (GSC) which utilizes a maximum likelihood technique for speaker localization. In the proposed implementation the adaptive GSC provides an estimate of the target signal as well as an estimation of target presence probability (TPP). Binaural outputs are generated in two different ways: In the first approach the binaural signals are rendered using the GSC output signal combined with the HRTF hypotheses which are adapted by the broadband localization. The second approach uses the GSC output and the TPP to determine a common spectral postfilter. We find that the adaptive beamformer combined with the binaural rendering technique leads to larger improvements of the quality of the desired signal and delivers less unnatural fluctuations as compared to the common spectral postfilter. Informal subjective tests as well as instrumental measurements in the presence of the listener head movements reveals, however, the benefit of the spatially motivated spectral postfilter for the preservation of binaural cues of both target and interferer signals. Mehdi Zohourian, Rainer Martin 0001 |
ICASSP | 2 |
| 2018 | A Noise Reduction Postfilter for Binaurally Linked Single-Microphone Hearing Aids Utilizing a Nearby External MicrophoneabstractThe use of a nearby external microphone for addressing front-back ambiguity in single-microphone hearing aid devices is investigated. Strategic placement of the external microphone is able to provide benefits from the body shielding back-directional noise and therefore information for discriminating between the frontal and back hemispheres. The scattering effects of the body are first analyzed to yield a placement strategy of the external microphone for maximizing the shielding effect of the body against back-directional noise sources while also optimizing for speech intelligibility. Assuming optimal placement, a frontal target source presence probability (FTSPP) estimator is derived. Using the FTSPP estimator, a more comprehensive noise estimator is proposed, which considers both stationary and nonstationary interferers. The performance of the proposed noise estimator is evaluated in its application for postfiltering the output of a binaural beamformer. The effect of postfiltering using the proposed noise estimator is to provide further reduction of directional noise from the lateral and back direction while preserving the frontal target signal. The resulting enhancement with the proposed noise estimator provides a better signal-to-noise ratio and improved objective speech quality and intelligibility of a frontal target speaker compared to the state of the art. Dianna Yee, A. Homayoun Kamkar-Parsi, Rainer Martin 0001, Henning Puder |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Binaural Speaker Localization Integrated Into an Adaptive Beamformer for Hearing AidsabstractIn this paper, we present and compare novel algorithms to localize simultaneous speakers using four microphones distributed on a pair of binaural hearing aids. The framework consists of two groups of localization algorithms, namely, beamforming-based and statistical model based localization algorithms. We first generalize our previously proposed methods based on beamforming techniques to the binaural configuration with 2 × 2 microphones. Next, we contribute two statistical model based methods for binaural localization using the maximum likelihood approach that also takes head-related transfer functions and unknown noise conditions into account. The methods enable the localization of multiple source positions for all azimuth angles and do not require prior training of binaural cues. The proposed localization algorithms are integrated into a generalized side-lobe canceller (GSC) to extract the desired speaker in the presence of competing speakers and background noise and when the head of the listener turns. The GSC components are adapted with the frequency-wise target presence probability and the frame-wise broadband direction-of-arrival (DOA) estimates that track the turns of the listener's head. We evaluate the performance of the localization algorithms individually and also in the context of the adaptive binaural beamformer in various noisy and reverberant conditions. Finally, we introduce a new adaptive beamformer, which combines the GSC with multichannel speech presence probability estimation and achieves superior source separation performance in noisy environment. Mehdi Zohourian, Gerald Enzner, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2018 | Spoofing Detection in Automatic Speaker Verification Systems Using DNN Classifiers and Dynamic Acoustic FeaturesabstractWith the development of speech synthesis technology, automatic speaker verification (ASV) systems have encountered the serious challenge of spoofing attacks. In order to improve the security of ASV systems, many antispoofing countermeasures have been developed. In the front-end domain, much research has been conducted on finding effective features which can distinguish spoofed speech from genuine speech and the published results show that dynamic acoustic features work more effectively than static ones. In the back-end domain, Gaussian mixture model (GMM) and deep neural networks (DNNs) are the two most popular types of classifiers used for spoofing detection. The log-likelihood ratios (LLRs) generated by the difference of human and spoofing log-likelihoods are used as spoofing detection scores. In this paper, we train a five-layer DNN spoofing detection classifier using dynamic acoustic features and propose a novel, simple scoring method only using human log-likelihoods (HLLs) for spoofing detection. We mathematically prove that the new HLL scoring method is more suitable for the spoofing detection task than the classical LLR scoring method, especially when the spoofing speech is very similar to the human speech. We extensively investigate the performance of five different dynamic filter bank-based cepstral features and constant Q cepstral coefficients (CQCC) in conjunction with the DNN-HLL method. The experimental results show that, compared to the GMM-LLR method, the DNN-HLL method is able to significantly improve the spoofing detection accuracy. Compared with the CQCC-based GMM-LLR baseline, the proposed DNN-HLL model reduces the average equal error rate of all attack types to 0.045%, thus exceeding the performance of previously published approaches for the ASVspoof 2015 Challenge task. Fusing the CQCC-based DNN-HLL spoofing detection system with ASV systems, the false acceptance rate on spoofing attacks can be reduced significantly. Hong Yu 0006, Zheng-Hua Tan, Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | Segmentation of music signals based on explained variance ratio for applications in spectral complexity reductionabstractSince natural acoustic signals like speech or music exhibit a highly varying temporal structure, signal enhancement and feature extraction algorithms benefit from segmentation procedures which take the underlying signal structure into account. In this paper we present a novel unsupervised segmentation procedure for music signals which relies on an explained variance criterion in the eigenspace of the constant-Q spectral domain. The procedure is used in the context of a spectral complexity reduction method which mitigates effects of cochlear hearing loss. It is compared to a segmentation based on equidistant boundaries. The results demonstrate that the proposed segmentation procedure gives an improvement in terms of signal-to-artefacts ratio in comparison to corresponding equidistant boundaries segmentation. Ekaterina A. Krymova, Anil M. Nagathil, Denis Belomestny, Rainer Martin 0001 |
ICASSP | 4 |
| 2017 | A feature-based linear regression model for predicting perceptual ratings of music by cochlear implant listenersabstractWhile speech quality and intelligibility prediction methods for normal-hearing and hearing-impaired listeners have found a lot of attention as a cost-saving complement to listening tests, analogous procedures for music signals are still rare. In this paper a method is proposed for predicting perceptual ratings of music as obtained by cochlear implant (CI) listeners. For this purpose a listening test with CI listeners was conducted, who were asked to provide their ratings for music excerpts on different scales. It is shown that principal component regression (PCR) is a suitable tool to model and accurately predict the median ratings of the CI listeners using timbre and pitch related signal features as predictor variables. These features describe signal characteristics such as high-frequency energy, spectral bandwidth and roughness. The proposed prediction model is a first step towards an instrumental evaluation procedure for music processing algorithms in hearing devices. Anil M. Nagathil, Jan-Willem Schlattmann, Katrin Neumann, Rainer Martin 0001 |
ICASSP | 4 |
| 2017 | Analysis of temporal aggregation and dimensionality reduction on feature sets for speaker identification in wireless acoustic sensor networksabstractIn this paper we analyze the impact of temporal feature aggregation and feature dimensionality reduction on the performance of speaker identification tasks. We investigate these two processing steps in the context of communication layer constraints, such as limited bitrate, and privacy constraints at node level, of a wireless acoustic sensor network. To this end, we extract Modulation-MFCC features and state-of-the-art i-vectors for speaker identification, and investigate temporal aggregation and dimensionality reduction in the feature extraction process. In the evaluation, we use clean data as well as reverberant data to assess the feature sets for different application scenarios. It is found that temporal aggregation has a positive effect on increasing speaker identification performance while respecting the aforementioned privacy constraints and that Linear Discriminant Analysis can be successfully employed for dimensionality reduction. Alexandru Nelus, Sebastian Gergen, Rainer Martin 0001 |
MMSP | 3 |
| 2016 | Efficient estimation of inter-subband speech correlationsabstractWe propose an approach to compute the inter-subband correlation (ISBC) of noisy speech signals to distinguish between speech and noise segments in the time-frequency plane. The proposed spectral correlation estimator provides information about the input signal which can be used to derive a binary mask or the speech-presence probability. Unlike other approaches it does not require an estimate of the noise power. To this end we analyse a received noisy speech signal in the modulation domain and identify similarly modulated subband signals within a range of modulation frequencies that are typical for speech signals. Based on this pre-processing step, we identify a single reference subband that most likely contains speech and estimate the spectral correlations with respect to this reference band. The algorithm proposed in this paper aims at a very low computational complexity which makes it suitable for hearing aids. Alexander Schasse, Rainer Martin 0001, Ulrich Kornagel, Eghart Fischer, Henning Puder |
ICASSP | 2 |
| 2016 | A speech enhancement system using binaural hearing aids and an external microphoneabstractThis paper presents a strategy for using an external microphone for enhancing noisy speech in single-microphone completely-in-canal (CIC) hearing aids. The external microphone is placed such that it benefits from the body shielding noise from the back hemisphere. The presented algorithm first enhances the external microphone signal without assuming an exact known location of the external microphone. The proposed algorithm then automatically incorporates the enhanced external microphone signal for post-processing enhancement of a conventional dual-channel binaural beamformer whenever the external microphone has a significant SNR advantage. The overall enhancement scheme avoids error-prone estimations of target voice activity detection and relative transfer functions between the microphones. Unlike single-channel post-processing filters which are limited to reducing stationary or diffuse noise, the proposed postprocessing filter is able to reduce highly non-stationary directional noise from the back hemisphere. The resulting system provides enhancement even for noise arriving from the backward direction, which conventionally is difficult for CIC hearing aids where there exists a front-back ambiguity. Dianna Yee, A. Homayoun Kamkar-Parsi, Henning Puder, Rainer Martin 0001 |
ICASSP | 4 |
| 2016 | Binaural speaker localization and separation based on a joint ITD/ILD model and head movement trackingabstractIn this paper we present a novel algorithm to localize and separate simultaneous speakers using hearing aids when the head is subject to rotational movement. Most of the algorithms used in hearing aids are able to extract target signals that are in the look direction of the user and suffer from a reduced performance in localizing sounds received from other directions. Moreover, head-shadowing as well as variations like head movements may lead to significant distortions. The proposed binaural GSC beamformer includes an MMSE-based localization algorithm using an ITD/ILD model and is controlled by an inertial measurement unit. The localization algorithm can effectively localize multiple speakers in the presence of reverberation. The estimated source locations are used to adapt the GSC beamformer which extracts the desired speaker. Experimental results demonstrate the performance of the new system and especially the benefits of ILD information. Mehdi Zohourian, Rainer Martin 0001 |
ICASSP | 2 |
| 2016 | Spectral Complexity Reduction of Music Signals for Mitigating Effects of Cochlear Hearing LossabstractIn this paper we study reduced-rank approximations of music signals in the constant-Q spectral domain as a means to reduce effects stemming from cochlear hearing loss. The rationale behind computing reduced-rank approximations is that they allow to reduce the spectral complexity of a music signal. The method is motivated by studies with cochlear implant listeners which have shown that solo instrumental music or music remixed at higher signal-to-interference ratios are preferred over complex music ensembles or orchestras. For computing the reduced-rank approximations we investigate methods based on principal component analysis and partial least squares analysis, and compare them to source separation algorithms. The strategies, which are applied to music with a predominant leading voice, are compared in terms of their ability for mitigating effects of simulated reduced frequency selectivity and with respect to source signal distortions. Established instrumental measures and a newly developed measure indicate a considerable reduction of the auditory distortion resulting from cochlear hearing loss. Furthermore, a listening test reveals a significant preference for the reduced-rank approximations in terms of melody clarity and ease of listening. Anil M. Nagathil, Claus Weihs, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2016 | A Frequency-Domain Adaptive Line Enhancer With Step-Size Control Based on Mutual Information for Harmonic Noise ReductionabstractWe propose an adaptive line enhancer with a frequency-dependent step-size. The proposed frequency-domain adaptive line enhancer is used as a single-channel noise reduction system for removing harmonic noise from noisy speech. Our main contribution is to exploit the temporal dependence in the log-magnitude and phase spectra of the noisy speech using mutual information, and to derive a frequency-dependent step-size which detects the presence of harmonic noise in different frequency bins. Our proposed step-size control allows the suppression of harmonic noise and the preservation of speech components. The experiments are performed with different real-life acoustic noises which contain harmonic components. Using instrumental speech intelligibility and quality measures, we demonstrate that the proposed approach can outperform the conventional frequency-domain adaptive line enhancer with a fixed step-size for harmonic noise reduction. Jalal Taghia, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2015 | Binaural speech enhancement with instantaneous coherence smoothing using the cepstral correlation coefficientabstractIn this paper we propose a novel approach to cepstral smoothing for reducing musical noise fluctuations in binaural speech enhancement. Similar to other methods, our approach computes a preliminary spectral gain function using the magnitude-squared coherence function and applies an instantaneous weighting to the gain function in the cepstral domain. In this contribution, the weighting function is based on the binaural cepstral correlation coefficient (CCC). We introduce the CCC and briefly discuss its properties. Similar to the cepstrum, the CCC emphasizes the spectral envelope and fundamental frequency information of the target signal, however, in a representation normalized to a range of [-1,1] and less sensitive to spatially uncorrelated noise. Thus, it can be easily and effectively used as a weight in the cepstral domain. The utility of the CCC is confirmed via experiments with different noise types and several instrumental measures. Rainer Martin 0001, Masoumeh Azarpour, Gerald Enzner |
ICASSP | 1 |
| 2015 | Multi-channel speaker localization and separation using a model-based GSC and an inertial measurement unitabstractIn this paper we propose a novel multi-channel algorithm to separate simultaneous speakers in an environment where the microphone array is subject to movement. When the microphones are mounted to a person's head, for instance, the movements can lead to ambiguities with respect to the sources and to distortions in the processed signal. The proposed system estimates the direction-of-arrival of the speaker's signals relative to the array and updates these estimates using an inertial measurement unit (IMU). A GMM-based localization model is used to compute the posterior probabilities of source activity in each time-frequency bin and its parameters are re-estimated during array movements. Then, a model-based generalized side-lobe canceler (GSC) whose components are continuously updated, is employed for the separation of sources. For various speeds of microphone array rotation, it is demonstrated that the IMU-based system delivers improved speech quality when compared to the baseline technique without IMU. Mehdi Zohourian, Alan Archer-Boyd, Rainer Martin 0001 |
ICASSP | 3 |
| 2015 | Reduction of reverberation effects in the MFCC modulation spectrum for improved classification of acoustic signals
Sebastian Gergen, Anil M. Nagathil, Rainer Martin 0001 |
INTERSPEECH | 3 |
| 2015 | Classification of reverberant audio signals using clustered ad hoc distributed microphones
Sebastian Gergen, Anil M. Nagathil, Rainer Martin 0001 |
Signal Process. | 3 |
| 2015 | Two-Stage Filter-Bank System for Improved Single-Channel Noise Reduction in Hearing AidsabstractThe filter-bank system implemented in hearing aids has to fulfill various constraints such as low latency and high stop-band attenuation, usually at the cost of low frequency resolution. In the context of frequency-domain noise-reduction algorithms, insufficient frequency resolution may lead to annoying residual noise artifacts since the spectral harmonics of the speech cannot properly be resolved. Especially in case of female speech signals, the noise between the spectral harmonics causes a distinct roughness of the processed signals. Therefore, this work proposes a two-stage filter-bank system, such that the frequency resolution can be improved for the purpose of noise reduction, while the original first-stage hearing-aid filter-bank system can still be used for compression and amplification. We also propose methods to implement the second filter-bank stage with little additional algorithmic delay. Furthermore, the computational complexity is an important design criterion. This finally leads to an application of the second filter-bank stage to lower frequency bands only, resulting in the ability to resolve the harmonics of speech. The paper presents a systematic description of the second filter-bank stage, discusses its influence on the processed signals in detail and further presents the results of a listening test which indicates the improved performance compared to the original single-stage filter-bank system. Alexander Schasse, Timo Gerkmann, Rainer Martin 0001, Wolfgang Sörgel, Thomas Pilgrim, Henning Puder |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Binaural noise PSD estimation for binaural speech enhancementabstractIn this paper we propose a novel binaural algorithm to estimate the power spectral density (PSD) of the background noise at the left and right ear separately. Inspired by equalization-cancelation considered in binaural hearing, the target speech is canceled at both left and right ears by means of the FLMS (fast least-mean square) algorithm. Assuming the ideal equalization, the output error of the blocking filter is a biased estimation of noise PSD. The estimated noise PSD is further corrected by exploiting the estimated left and right interaural transfer functions and the coherence model of the noise field. In addition to noise power estimation assessment, the estimated noise PSD is integrated in a binaural speech enhancement framework in order to evaluate the overall noise reduction performance. Masoumeh Azarpour, Gerald Enzner, Rainer Martin 0001 |
ICASSP | 3 |
| 2014 | Nonlinear estimation of missing ΔLSF parameters by a mixture of Dirichlet distributionsabstractIn packet networks, a reliable scheme to handle packet loss during speech transmission is of great importance. As a common representation of the linear predictive coding (LPC) model, the line spectral frequency (LSF) parameters are widely used in speech quantization and transmission. In this paper, we propose a novel scheme to estimate the missing values occurring during LPC model transmission. In order to exploit the boundary and ordering properties of the LSF parameters, we utilize the ΔLSF representation and apply the Dirichlet mixture model (DMM) to capture the correlations among the elements in the ΔLSF vector. With the conditional distribution of the missing part given the received part, an optimal nonlinear minimum mean square error estimator for the missing values is proposed. Compared to the previously presented Gaussian mixture model based method, the proposed DMM based nonlinear estimator shows a convincing improvement. Zhanyu Ma, Rainer Martin 0001, Jun Guo 0002, Honggang Zhang 0002 |
ICASSP | 2 |
| 2014 | A negentropy based adaptive line enhancer for single-channel noise reduction at low SNR conditionsabstractIn this paper, we propose an adaptive line enhancer based on negentropy for single-channel noise reduction. Our proposed approach can be integrated in a speech enhancement system as a preprocessor to be combined with other noise reduction approaches. The proposed method performs the noise reduction by splitting the noisy speech components into the deterministic and the stochastic parts through the minimization of negentropy in an adaptive manner. We consider the negentropy as a cost function, and we derive a learning rule via Newton's method to minimize the negentropy of the error signal. By the experimental results, we demonstrate that exploiting the proposed approach can be potentially useful as a preprocessor for improving the performance of conventional single-channel noise reduction approaches at low signal-to-noise ratio (SNR) conditions. Moreover, it is shown that our approach by itself can also enhance the noisy speech in an adverse noisy environment. Jalal Taghia, Rainer Martin 0001 |
ICASSP | 2 |
| 2014 | Scalable low-complexity GPS and DGPS positioning using approximate QR decomposition
Yuheng He, Rainer Martin 0001, Attila Michael Bilgic |
Signal Process. | 2 |
| 2014 | Estimation of Subband Speech Correlations for Noise Reduction via MVDR ProcessingabstractRecently, it has been proposed to use the minimum-variance distortionless-response (MVDR) approach in single-channel speech enhancement in the short-time frequency domain. By applying optimal FIR filters to each subband signal, these filters reduce additive noise components with less speech distortion compared to conventional approaches. An important ingredient to these filters is the temporal correlation of the speech signals. We derive algorithms to provide a blind estimation of this quantity based on a maximum-likelihood and maximum a-posteriori estimation. To derive proper models for the inter-frame correlation of the speech and noise signals, we investigate their statistics on a large dataset. If the speech correlation is properly estimated, the previously derived subband filters discussed in this work show significantly less speech distortion compared to conventional noise reduction algorithms. Therefore, the focus of the experimental parts of this work lies on the quality and intelligibility of the processed signals. To evaluate the performance of the subband filters in combination with the clean speech inter-frame correlation estimators, we predict the speech quality and intelligibility by objective measures. Alexander Schasse, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2014 | Variational Bayesian Inference for Multichannel Dereverberation and Noise ReductionabstractRoom reverberation and background noise severely degrade the quality of hands-free speech communication systems. In this work, we address the problem of combined speech dereverberation and noise reduction using a variational Bayesian (VB) inference approach. Our method relies on a multichannel state-space model for the acoustic channels that combines frame-based observation equations in the frequency domain with a first-order Markov model to describe the time-varying nature of the room impulse responses. By modeling the channels and the source signal as latent random variables, we formulate a lower bound on the log-likelihood function of the model parameters given the observed microphone signals and iteratively maximize it using an online expectation-maximization approach. Our derivation yields update equations to jointly estimate the channel and source posterior distributions and the remaining model parameters. An inspection of the resulting VB algorithm for blind equalization and channel identification (VB-BENCH) reveals that the presented framework includes previously proposed methods as special cases. Finally, we evaluate the performance of our approach in terms of speech quality, adaptation times, and speech recognition results to demonstrate its effectiveness for a wide range of reverberation and noise conditions. Dominic Schmid, Gerald Enzner, Sarmad Malik, Dorothea Kolossa, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2014 | Objective Intelligibility Measures Based on Mutual Information for Speech Subjected to Speech Enhancement ProcessingabstractWe propose a novel method for objective speech intelligibility prediction which can be useful in many application domains such as hearing instruments and forensics. Most objective intelligibility measures available in the literature employ some kind of signal-to-noise ratio (SNR) or a correlation-based comparison between the spectro-temporal representations of clean and processed speech. In this paper, we investigate the speech intelligibility prediction from the viewpoint of information theory and introduce novel objective intelligibility measures based on the estimated mutual information between the temporal envelopes of clean speech and processed speech in the subband domain. Mutual information allows to account for higher order statistics and hence to consider dependencies beyond the conventional second order statistics. Using data from three different listening tests it is shown that the proposed objective intelligibility measures provide promising results for speech intelligibility prediction in different scenarios of speech enhancement where speech is processed by non-linear modification strategies. Jalal Taghia, Rainer Martin 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Performance of Specific vs. Generic Feature Sets in Polyphonic Music Instrument Recognition
Igor Vatolkin, Anil M. Nagathil, Wolfgang M. Theimer, Rainer Martin 0001 |
EMO | 4 |
| 2013 | Adaptive binaural noise reduction based on matched-filter equalization and post-filteringabstractIn this paper a binaural noise reduction system based on the adaptive matched filter array (MFA) and post-filtering is presented. The binaural MFA filter is formed using the interaural impulse response between left and right microphone signal which is estimated by means of the NLMS algorithm. The residual noise reduction by three post-filter types is then compared and evaluated according to objective and subjective measures. Furthermore, the performance of the algorithm for preserving binaural cues is discussed. Masoumeh Azarpour, Gerald Enzner, Rainer Martin 0001 |
ICASSP | 3 |
| 2013 | Improved single-channel nonstationary noise tracking by an optimized MAP-based postprocessorabstractIn this paper we present an improved version of the recently proposed Maximum A-Posteriori (MAP) based noise power spectral density estimator. An empirical bias compensation and bandwidth adjustment reduce bias and variance of the noise variance estimates. The main advantage of the MAP-based postprocessor is its low estimation variance. The estimator is employed in the second stage of a two-stage single-channel speech enhancement system, where eight different state-of-the-art noise tracking algorithms were tested in the first stage. While the postprocessor hardly affects the results in stationary noise scenarios, it becomes the more effective the more nonstationary the noise is. The proposed postprocessor was able to improve all systems in babble noise w.r.t. the perceptual evaluation of speech quality performance. Aleksej Chinaev, Reinhold Häb-Umbach, Jalal Taghia, Rainer Martin 0001 |
ICASSP | 4 |
| 2013 | Audio signal classification in reverberant environments based on fuzzy-clustered ad-hoc microphone arraysabstractAudio signal classification suffers from the mismatch of environmental conditions when training data is based on clean and anechoic signals and test data is distorted by reverberation and signals from other sources. In this contribution we analyze the classification performance for such a scenario with two concurrently active sources in a simulated reverberant environment. To obtain robust classification results, we exploit the spatial distribution of ad-hoc microphone arrays to capture the signals and extract cepstral features. Based on these features only, we use unsupervised fuzzy clustering to estimate clusters of microphones which are dominated by one of the sources. Finally, signal classification based on clean and anechoic training data is performed for each of the cluster. The probability of cluster membership for each microphone is provided by the fuzzy clustering algorithm and is used to compute a weighted average of the feature vectors. It is shown that the proposed method exceeds the performance of classification based on single microphones. Sebastian Gergen, Anil M. Nagathil, Rainer Martin 0001 |
ICASSP | 3 |
| 2013 | Online inter-frame correlation estimation methods for speech enhancement in frequency subbandsabstractIn this paper, we propose solutions for the online adaptation of optimal FIR filters for speech enhancement in DFT subbands. An important ingredient to such filters is the estimation of the inter-frame correlation of the clean speech signal. While this correlation was assumed to be perfectly known in former studies, we discuss two online estimation approaches based on a constant noise inter-frame correlation and on the use of a binary mask. We show that a filtering of subband signals based on these estimated quantities outperforms a conventional, instantaneous spectral weighting, such as the frequency-domain Wiener filter at least for high SNR conditions. Alexander Schasse, Rainer Martin 0001 |
ICASSP | 2 |
| 2013 | Dual-channel noise reduction based on a mixture of circular-symmetric complex Gaussians on unit hypersphereabstractIn this paper a model-based dual-channel noise reduction approach is presented which is an alternative to conventional noise reduction algorithms essentially due to its independence of the noise power spectral density estimation and of any prior knowledge about the spatial noise field characteristics. We use a mixture of circular-symmetric complex-Gaussian distributions projected on the unit hypersphere for modeling the complex discrete Fourier transform coefficients of noisy speech signals in the frequency domain. According to the derived mixture model, clustering of the noise and the target speech components is performed depending on their direction of arrival. A soft masking strategy is proposed for speech enhancement based on responsibilities assigned to the target speech class in each time-frequency bin. Our experimental results show that the proposed approach is more robust than conventional dual-channel noise reduction systems based on the single- and dual-channel noise power spectral density estimators. Jalal Taghia, Rainer Martin 0001, Jalil Taghia, Arne Leijon |
ICASSP | 2 |
| 2013 | Integration of beamforming and uncertainty-of-observation techniques for robust ASR in multi-source environments
Ramón Fernandez Astudillo, Dorothea Kolossa, Alberto Abad, Steffen Zeiler, Rahim Saeidi, Pejman Mowlaee, João Paulo da Silva Neto, Rainer Martin 0001 |
Comput. Speech Lang. | 8 |
| 2013 | Spectral Domain Speech Enhancement Using HMM State-Dependent Super-Gaussian PriorsabstractThe derivation of MMSE estimators for the DFT coefficients of speech signals, given an observed noisy signal and super-Gaussian prior distributions, has received a lot of interest recently. In this letter, we look at the distribution of the periodogram coefficients of different phonemes, and show that they have a gamma distribution with shape parameters less than one. This verifies that the DFT coefficients for not only the whole speech signal but also for individual phonemes have super-Gaussian distributions. We develop a spectral domain speech enhancement algorithm, and derive hidden Markov model (HMM) based MMSE estimators for speech periodogram coefficients under this gamma assumption in both a high uniform resolution and a reduced-resolution Mel domain. The simulations show that the performance is improved using a gamma distribution compared to the exponential case. Moreover, we show that, even though beneficial in some aspects, the Mel-domain processing does not lead to better results than the algorithms in the high-resolution domain. Nasser Mohammadiha, Rainer Martin 0001, Arne Leijon |
IEEE Signal Process. Lett. | 2 |
| 2013 | Corpus-Based Speech Enhancement With Uncertainty Modeling and Cepstral SmoothingabstractWe present a new approach for corpus-based speech enhancement that significantly improves over a method published by Xiao and Nickel in 2010. Corpus-based enhancement systems do not merely filter an incoming noisy signal, but resynthesize its speech content via an inventory of pre-recorded clean signals. The goal of the procedure is to perceptually improve the sound of speech signals in background noise. The proposed new method modifies Xiao's method in four significant ways. Firstly, it employs a Gaussian mixture model (GMM) instead of a vector quantizer in the phoneme recognition front-end. Secondly, the state decoding of the recognition stage is supported with an uncertainty modeling technique. With the GMM and the uncertainty modeling it is possible to eliminate the need for noise dependent system training. Thirdly, the post-processing of the original method via sinusoidal modeling is replaced with a powerful cepstral smoothing operation. And lastly, due to the improvements of these modifications, it is possible to extend the operational bandwidth of the procedure from 4 kHz to 8 kHz. The performance of the proposed method was evaluated across different noise types and different signal-to-noise ratios. The new method was able to significantly outperform traditional methods, including the one by Xiao and Nickel, in terms of PESQ scores and other objective quality measures. Results of subjective CMOS tests over a smaller set of test samples support our claims. Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 4 |
| 2012 | On the construction of window functions with constant-overlap-add constraint for arbitrary window shiftsabstractIn this paper we present a construction method for window functions with constant-overlap-add (COLA) constraint for spectral analysis-synthesis with a given percentage of window overlap. The window functions are derived from the shortest possible COLA window i.e., the rectangular window. The construction method allows to adjust the side-lobe fall-off and to fulfill the COLA constraint for arbitrary window shifts. The procedure results in a family of window functions which includes some well known members (the Bartlett and the Hann window) for 50% overlap. Christian Borß, Rainer Martin 0001 |
ICASSP | 2 |
| 2012 | Subjective and objective quality assessment of single-channel speech separation algorithmsabstractPrevious studies on performance evaluation of single-channel speech separation (SCSS) algorithms mostly focused on automatic speech recognition (ASR) accuracy as their performance measure. Assessing the separated signals by different metrics other than this has the benefit that the results are expected to carry on to other applications beyond ASR. In this paper, in addition to conventional speech quality metrics (PESQ and SNRloss), we also evaluate the separation systems output using different source separation metrics: blind source separation evaluation (BSS EVAL) and perceptual evaluation methods for audio source separation (PEASS) measures. In our experiments, we apply these measures on the separated signals obtained by two well-known systems in the SCSS challenge to assess the objective and subjective quality of their output signals. Comparing subjective and objective measurements shows that PESQ and PEASS quality metrics predict well the subjective quality of separated signals obtained by the separation systems. From the results it is observed that the short-time objective intelligibility (STOI) measure predict the speech intelligibility results. Pejman Mowlaee, Rahim Saeidi, Mads Græsbøll Christensen, Rainer Martin 0001 |
ICASSP | 4 |
| 2012 | Optimal signal reconstruction from a constant-Q spectrumabstractIn contrast to other well-known techniques for spectral analysis such as the discrete Fourier transform or the wavelet transform the constant-Q transform (CQT) matches the center frequencies of its sub-band filters to the frequency scale of western music and accounts for the requirement of frequency-dependent bandwidths. However, it does not possess a strict mathematical inverse. Therefore, we derive an optimal reconstruction method to recover a signal from its CQT spectrum. The reconstruction problem is posed as a segmented overdetermined minimization problem where the number of frequency bins is larger than the number of reconstructed signal samples in each segment. The framework also enables the reduction of the input-output latency of the CQT. The method is evaluated for different types of music signals as well as for speech and Gaussian noise. For classical music a reconstruction quality of 89 dB signal-to-noise ratio is achieved. Anil M. Nagathil, Rainer Martin 0001 |
ICASSP | 2 |
| 2012 | Inventory-style speech enhancement with uncertainty-of-observation techniquesabstractWe present a new method for inventory-style speech enhancement that significantly improves over earlier approaches [1]. Inventory-style enhancement attempts to resynthesize a clean speech signal from a noisy signal via corpus-based speech synthesis. The advantage of such an approach is that one is not bound to trade noise suppression against signal distortion in the same way that most traditional methods do. A significant improvement in perceptual quality is typically the result. Disadvantages of this new approach, however, include speaker dependency, increased processing delays, and the necessity of substantial system training. Earlier published methods relied on a-priori knowledge of the expected noise type during the training process [1]. In this paper we present a new method that exploits uncertainty-of-observation techniques to circumvent the need for noise specific training. Experimental results show that the new method is not only able to match, but outperform the earlier approaches in perceptual quality. Robert M. Nickel, Ramón Fernandez Astudillo, Dorothea Kolossa, Steffen Zeiler, Rainer Martin 0001 |
ICASSP | 5 |
| 2012 | On mutual information as a measure of speech intelligibilityabstractSpeech intelligibility prediction of noisy and processed noisy speech is important in a number of application domains such as hearing instruments and forensics. Most available objective intelligibility measures employ either a signal-to-noise ratio (SNR)-based or correlation-based comparison between frequency bands of the clean and the processed speech. In this paper, we approach the speech intelligibility prediction from the angle of information theory and show that an information theoretic concept provides a unified viewpoint on both the SNR and the correlation based approaches. Two objective intelligibility measures are introduced based on estimated mutual information between the clean speech and the processed speech in the time and the frequency subband domain. Our proposed measures show high correlation with subjective intelligibility measure (i.e. word correct scores) and comparative results with the short-term objective intelligibility measure (STOI). Jalal Taghia, Rainer Martin 0001, Richard C. Hendriks |
ICASSP | 2 |
| 2012 | Inventory-Based Audio-Visual Speech Enhancement
Dorothea Kolossa, Robert M. Nickel, Steffen Zeiler, Rainer Martin 0001 |
INTERSPEECH | 4 |
| 2012 | Phase estimation for signal reconstruction in single-channel source separationabstractSingle-channel speech separation algorithms frequently ignore the issue of accurate phase estimation while reconstructing the enhanced signal. Instead, they directly employ the mixed-signal phase for signal reconstruction which leads to undesired traces of the interfering source in the target signal. In this paper, assuming a given knowledge of signal spectrum amplitude, we present a solution to estimate the phase information for signal reconstruction of the sources from a single-channel mixture observation. We first investigate the effectiveness of the proposed phase estimation method employing known magnitude spectra of sources as an ideal case. We further relax the ideal signal spectra assumption by perturbing the clean signal spectra via Gaussian noise. The results show that for both scenarios, ideal and noisy magnitude signal spectra, the proposed phase estimation approach offers improved signal reconstruction accuracy, segmental SNR and PESQ compared to benchmark methods, and those neglecting the phase information. Pejman Mowlaee, Rahim Saeidi, Rainer Martin 0001 |
INTERSPEECH | 3 |
| 2011 | A model-based auditory scene analysis approach and its application to speech source localizationabstractLocalization and classification of acoustic signals in a complex auditory scene is an every day task of the human auditory system. However, this problem presents a significant challenge for a computational system. In this paper, we propose a framework that allows the detection of different classes of signals (e.g. speech, music) and the localization of the signal sources. As a special application of this framework, we describe a method for localization of human speech in a complex acoustic scene and low SNR, based on vowel detection. In contrast to other works, we detect and extract parts of the signal mixture that correspond to a single class exclusively, using a prior model of the target signals. Herewith, a reliable localization of signal sources is possible. Václav Bouse, Rainer Martin 0001 |
ICASSP | 2 |
| 2011 | Hierarchical audio classification using cepstral modulation ratio regressions based on Legendre polynomialsabstractIn this work we present a scalable feature set which is obtained by fitting orthogonal polynomials to the normalized modulation spectrum of cepstral coefficients and which can be easily adapted to different classification tasks. The performance of the feature set is investigated in a hierarchically structured audio signal classification experiment and compared with other approaches reported in the literature. For the root categories speech, music and noise a classification accuracy of 95% is achieved. Subclasses such as male and female speech or different noise types are classified with an accuracy of 95% and 85%, respectively. In a 10-category musical genre discrimination experiment the proposed features exhibit an accuracy of 61%. Anil M. Nagathil, Peter Gottel, Rainer Martin 0001 |
ICASSP | 3 |
| 2011 | An evaluation of noise power spectral density estimation algorithms in adverse acoustic environmentsabstractNoise power spectral density estimation is an important component of speech enhancement systems due to its considerable effect on the quality and the intelligibility of the enhanced speech. Recently, many new algorithms have been proposed and significant progress in noise tracking has been made. In this paper, we present an evaluation framework for measuring the performance of some recently proposed and some well-known noise power spectral density estimators and compare their performance in adverse acoustic environments. In this investigation we do not only consider the performance in the mean of a spectral distance measure but also evaluate the variance of the estimators as the latter is related to undesirable fluctuations also known as musical noise. By providing a variety of different non-stationary noises, the robustness of noise estimators in adverse environments is examined. Jalal Taghia, Jalil Taghia, Nasser Mohammadiha, Jinqiu Sang, Václav Bouse, Rainer Martin 0001 |
ICASSP | 6 |
| 2011 | Analysis of the Decision-Directed SNR Estimator for Speech Enhancement With Respect to Low-SNR and Transient ConditionsabstractBecause of their many applications and their relative ease of implementation, single-channel speech enhancement algorithms have received much attention. As a consequence, a vast amount of publications on estimation procedures and their implementation in noise reduction systems exists. However, there has been little systematic research on the theoretic performance of such estimators. In this paper, we provide a systematic analysis of the performance of noise reduction algorithms in low signal-to-noise ratio (SNR) and transient conditions, where we consider approaches using the well-known decision-directed SNR estimator. We show that the smoothing properties of the decision-directed SNR estimator in low SNR conditions can be analytically described and that the limits of noise reduction for widely used spectral speech estimators based on the decision-directed approach can be predicted. We also illustrate that achieving both a good preservation of speech onsets in transient conditions on one side and the suppression of musical noise on the other can be especially problematic when the decision-directed SNR estimation is used. Colin Breithaupt, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | A Versatile Framework for Speaker Separation Using a Model-Based Speaker Localization ApproachabstractWe build upon our speaker localization framework developed in a previous work (N. Madhu and R. Martin, A scalable framework for multiple speaker localization and tracking,” in Proc. Int. Workshop Acoustic Echo Noise Control (IWAENC), Sep. 2008) to perform source separation. The proposed approach, exploiting the supplementary information from the mixture of Gaussians-based localization model, allows for the incorporation of a wide class of separation algorithms, from the nonlinear time-frequency mask-based approaches to a fully adaptive beamformer in the generalized sidelobe canceller (GSC) structure. We propose, in addition, a generalized estimation of the blocking matrix based on subspace projectors. The adaptive beamformer realized as proposed is insensitive to gain mismatches among the sensors, obviating the need for magnitude calibration of the microphones. It is also demonstrated that the proposed linear approach has a performance comparable to that of an optimal (oracle) GSC implementation. In comparison to ICA-based approaches, another advantage of the separation framework described herein is its robustness to ambient noise and scenarios with an unknown number of sources. Nilesh Madhu, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Speech presence probability estimation based on temporal cepstrum smoothingabstractWe propose a novel, robust estimator for the probability of speech presence at each time-frequency point in the short-time discrete Fourier domain. While existing estimators perform quite reliably in stationary noise environments, they usually exhibit a large false-alarm rate in nonstationary noise that results in a great deal of noise leakage when applied to a speech enhancement task. The proposed estimator overcomes this problem by temporally smoothing the cepstrum of the a posteriori signal-to-noise ratio (SNR), and yields considerably less noise leakage and low speech distortions in both, stationary and nonstationary noise as compared to state-of-the-art estimators. Especially in babble noise, this results in large SNR improvements. Timo Gerkmann, Martin Krawczyk-Becker, Rainer Martin 0001 |
ICASSP | 3 |
| 2010 | Optimization of switchable windows for low-delay spectral analysis-synthesisabstractWe present a novel iterative method for the optimization of switchable pairs of window functions. These windows may be used for block-based spectral analysis-synthesis (AS) in low-delay speech enhancement systems, where the energy compaction of speech sounds is improved by switching the spectral AS windows. Optimization objectives of the approach take the frequency response, quasi perfect reconstruction (PR) of each window pair and quasi-PR during window switching into account. An example of window pairs obtained with the proposed method clearly outperforms a reference design. The improved aliasing and imaging suppression is particularly important for hearing aids where high spectral gains may lead to audible reconstruction artifacts. Dirk Mauler, Rainer Martin 0001 |
ICASSP | 2 |
| 2009 | Multi-microphone maximum a posteriori fundamental frequency estimation in the cepstral domainabstractIn this work we derive a new cepstrum based maximum likelihood fundamental frequency estimator that exploits the information of multiple microphones. The new approach results in a maximum search on the sum of the microphone cepstra. We compare the new approach to a maximum search on the cepstrum of the output signal of a delay-and-sum beamformer. We show that the new approach outperforms the beamforming approach for all considered input signal-to-noise ratios. We develop a general framework which includes the cepstral harmonics of the fundamental frequency and extend the approach towards a maximum a posteriori fundamental period tracker that further enhances the results and increases the robustness in noisy environments. Timo Gerkmann, Rainer Martin 0001, Derya Dalga |
ICASSP | 2 |
| 2009 | Cepstral modulation ratio regression (CMRARE) parameters for audio signal analysis and classificationabstractIn this paper we propose a new set of parameters for audio signal analysis and classification. These parameters are regressions computed on the normalized modulation spectrum of high-resolution cepstral coefficients. The parameter set is scalable in its size and gives a compact representation of the modulation content of speech and other audio signals. These parameters as well as the regression approximation error are well suited for characterizing audio signals in a unified framework. In particular we use a set of eight parameters in a speech/music/noise classification task in which we achieve a classification accuracy which compares very well with other approaches including static and dynamic MFCCs. Rainer Martin 0001, Anil M. Nagathil |
ICASSP | 1 |
| 2008 | A novel a priori SNR estimation approach based on selective cepstro-temporal smoothingabstractWhile state-of-the-art approaches obtain an estimate of the a priori SNR by adaptively smoothing its maximum likelihood estimate in the frequency domain, we selectively smooth the maximum likelihood estimate in the cepstral domain. In the cepstral domain the noisy speech signal is decomposed into coefficients related mainly to the speech envelope, the excitation, and noise. As in the cepstral domain coefficients that represent speech can be robustly determined, we can apply little smoothing to speech coefficients and strong smoothing to noise coefficients. Thus, speech components are preserved and musical noise is suppressed. In speech enhancement experiments we obtain consistent improvements over the well known decision-directed approach. Colin Breithaupt, Timo Gerkmann, Rainer Martin 0001 |
ICASSP | 3 |
| 2008 | Parameterized MMSE spectral magnitude estimation for the enhancement of noisy speechabstractThe enhancement of short-term spectra of noisy speech can be achieved by statistical estimation of the clean speech spectral components. We present a minimum mean-square error estimator of the clean speech spectral magnitude that uses both a parametric compression function in the estimation error criterion and a parametric prior distribution for the statistical model of the clean speech magnitude. The novel parametric estimator has many known magnitude estimators as a special solution and, additionally, affords estimators that combine the beneficial properties of different known solutions. The new estimator is evaluated in terms of segmental SNR, speech distortion, and noise suppression. Colin Breithaupt, Martin Krawczyk-Becker, Rainer Martin 0001 |
ICASSP | 3 |
| 2008 | Temporal smoothing of spectral masks in the cepstral domain for speech separationabstractThis contribution details the development of a mask-based post- processor to improve the interference suppression in speech signals separated using linear deconvolution algorithms like independent component analysis (ICA). The design of the proposed post-filter is in two stages: in the first stage, use is made of the disjointness of the separated signals in the time-frequency domain to obtain binary masks to suppress cross-talk that generally remains after separation. In the next stage, a novel smoothing of the masks is proposed that preserves the speech structure of the target source while eliminating the random peaks in the time-frequency plane that lead to fluctuating background noise. The result is an enhanced signal with reduced cross-talk and no musical noise. Nilesh Madhu, Colin Breithaupt, Rainer Martin 0001 |
ICASSP | 3 |
| 2008 | Interactive auditory virtual environments for mobile devicesabstractIn this paper we present a client-server architecture which makes interactive auditory virtual environments (AVEs) accessible to devices with limited computational power. The implementation of an AVE generator as web service allows for platform independent "AVE services" via mobile devices almost "anywhere on any device" using a standard web browser. We propose a client-server architecture which computes the acoustic signals on a high-performance server and provides low-latency audio streaming from the server to the client. Christian Borß, Rainer Martin 0001 |
Mobile HCI | 2 |
| 2008 | Improved A Posteriori Speech Presence Probability Estimation Based on a Likelihood Ratio With Fixed PriorsabstractIn this paper, we present an improved estimator for the speech presence probability at each time-frequency point in the short-time Fourier transform domain. In contrast to existing approaches, this estimator does not rely on an adaptively estimated and thus signal-dependent a priori signal-to-noise ratio estimate. It therefore decouples the estimation of the speech presence probability from the estimation of the clean speech spectral coefficients in a speech enhancement task. Using both a fixed a priori signal-to-noise ratio and a fixed prior probability of speech presence, the proposed a posteriori speech presence probability estimator achieves probabilities close to zero for speech absence and probabilities close to one for speech presence. While state-of-the-art speech presence probability estimators use adaptive prior probabilities and signal-to-noise ratio estimates, we argue that these quantities should reflect true a priori information that shall not depend on the observed signal. We present a detection theoretic framework for determining the fixed a priori signal-to-noise ratio. The proposed estimator is conceptually simple and yields a better tradeoff between speech distortion and noise leakage than state-of-the-art estimators. Timo Gerkmann, Colin Breithaupt, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | On optimal estimation of compressed speech for hearing aids
Dirk Mauler, Anil M. Nagathil, Rainer Martin 0001 |
INTERSPEECH | 3 |
| 2007 | Cepstral Smoothing of Spectral Filter Gains for Speech Enhancement Without Musical NoiseabstractMany speech enhancement algorithms that modify short-term spectral magnitudes of the noisy signal by means of adaptive spectral gain functions are plagued by annoying spectral outliers. In this letter, we propose cepstral smoothing as a solution to this problem. We show that cepstral smoothing can effectively prevent spectral peaks of short duration that may be perceived as musical noise. At the same time, cepstral smoothing preserves speech onsets, plosives, and quasi-stationary narrowband structures like voiced speech. The proposed recursive temporal smoothing is applied to higher cepstral coefficients only, excluding those representing the pitch information. As the higher cepstral coefficients describe the finer spectral structure of the Fourier spectrum, smoothing them along time prevents single coefficients of the filter function from changing excessively and independently of their neighboring bins, thus suppressing musical noise. The proposed cepstral smoothing technique is very effective in nonstationary noise. Colin Breithaupt, Timo Gerkmann, Rainer Martin 0001 |
IEEE Signal Process. Lett. | 3 |
| 2007 | MAP Estimators for Speech Enhancement Under Normal and Rayleigh Inverse Gaussian DistributionsabstractThis paper presents a new class of estimators for speech enhancement in the discrete Fourier transform (DFT) domain, where we consider a multidimensional normal inverse Gaussian (MNIG) distribution for the speech DFT coefficients. The MNIG distribution can model a wide range of processes, from heavy-tailed to less heavy-tailed processes. Under the MNIG distribution complex DFT and amplitude estimators are derived. In contrast to other estimators, the suppression characteristics of the MNIG-based estimators can be adapted online to the underlying distribution of the speech DFT coefficients. Compared to noise suppression algorithms based on preselected super-Gaussian distributions, the MNIG-based complex DFT and amplitude estimators lead to a performance improvement in terms of segmental signal-to-noise ratio (SNR) in the order of 0.3 to 0.6 dB and 0.2 to 0.6 dB, respectively Richard C. Hendriks, Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Statistical analysis and performance of DFT domain noise reduction filters for robust speech recognition
Colin Breithaupt, Rainer Martin 0001 |
INTERSPEECH | 2 |
| 2006 | Soft decision combining for dual channel noise reduction
Timo Gerkmann, Rainer Martin 0001 |
INTERSPEECH | 2 |
| 2006 | Bias compensation methods for minimum statistics noise power spectral density estimation
Rainer Martin 0001 |
Signal Process. | 1 |
| 2005 | Robust speaker localization through adaptive weighted pair TDOA (AWEPAT) estimationabstractTime delay of arrival (TDOA) estimation between signals input to two or more microphones plays an important role in speaker localization.Most methods employ a linear array of two or more microphones and use the generalized cross correlation method or eigenspace analysis (AEDA) methods.TDOA estimation with linear arrays, however, is highly sensitive to estimation errors when the signals arrive from an endfire direction.In this paper we propose a novel adaptive algorithm which makes use of a three-microphone planar array.This algorithm exhibits a much smaller estimation error over the complete azimuth range of 0-360 degrees as compared to other algorithms.The computational complexity of this approach is comparable to other state-of-the-art algorithms. Nilesh Madhu, Rainer Martin 0001 |
INTERSPEECH | 2 |
| 2005 | Speech Enhancement Based on Minimum Mean-Square Error Estimation and Supergaussian PriorsabstractThis paper presents a class of minimum mean-square error (MMSE) estimators for enhancing short-time spectral coefficients of a noisy speech signal. In contrast to most of the presently used methods, we do not assume that the spectral coefficients of the noise or of the clean speech signal obey a (complex) Gaussian probability density. We derive analytical solutions to the problem of estimating discrete Fourier transform (DFT) coefficients in the MMSE sense when the prior probability density function of the clean speech DFT coefficients can be modeled by a complex Laplace or by a complex bilateral Gamma density. The probability density function of the noise DFT coefficients may be modeled either by a complex Gaussian or by a complex Laplacian density. Compared to algorithms based on the Gaussian assumption, such as the Wiener filter or the Ephraim and Malah (1984) MMSE short-time spectral amplitude estimator, the estimators based on these supergaussian densities deliver an improved signal-to-noise ratio. Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | MMSE estimation of magnitude-squared DFT coefficients with superGaussian priorsabstractWe present two minimum mean square error (MMSE) frequency domain estimators of the squared magnitude of a clean speech signal that is degraded by additive noise. These estimators are derived under the assumption that the DFT (discrete Fourier transform) coefficients of the clean speech are best modelled by the Gamma probability distribution function (PDF) instead of the common Gaussian PDF. The statistics of the perturbing noise is the Gaussian PDF in one case and the Laplacian PDF in the other. The estimators are used as noise reduction filters in the experimental evaluation. We give a comparison with a previously derived estimator which uses the Gaussian PDF as the PDF for speech and noise coefficients. Colin Breithaupt, Rainer Martin 0001 |
ICASSP (1) | 2 |
| 2002 | Unbiased residual echo power estimation for hands-free telephonyabstractResidual echo arises in hands-free telephony equipment due to insufficient echo canceler convergence, but can be suppressed using a postfilter. The most important control parameter for postfilter adaptation is therefore the residual echo power spectral density (PSD). In this contribution we present and compare residual echo PSD estimation techniques. We introduce a new partitioned block-adaptive estimator delivering unbiased residual echo PSD estimates in strongly reverberant and noisy acoustic environments. Gerald Enzner, Rainer Martin 0001, Peter Vary |
ICASSP | 2 |
| 2002 | Speech enhancement using MMSE short time spectral estimation with gamma distributed speech priorsabstractIn this paper we consider optimal estimators for speech enhancement in the Discrete Fourier Transform (DFT) domain. We present an analytical solution for estimating complex DFT coefficients in the MMSE sense when the clean speech DFT coefficients are Gamma distributed and the DFT coefficients of the noise are Gaussian or Laplace distributed. Compared to the state-of-the-art Wiener or MMSE short time amplitude estimators the new estimators deliver improved signal-to-noise ratios. When the noise model is a Laplacian density the enhanced speech shows less annoying random fluctuations in the residual noise than for a Gaussian density. Rainer Martin 0001 |
ICASSP | 1 |
| 2002 | A psychoacoustic approach to combined acoustic echo cancellation and noise reductionabstractThis paper presents and compares algorithms for combined acoustic echo cancellation and noise reduction for hands-free telephones. A structure is proposed, consisting of a conventional acoustic echo canceler and a frequency domain postfilter in the sending path of the hands-free system. The postfilter applies the spectral weighting technique and attenuates both the background noise and the residual echo which remains after imperfect echo cancellation. Two weighting rules for the postfilter are discussed. The first is a conventional one, known from noise reduction, which is extended to attenuate residual echo as well as noise. The second is a psychoacoustically motivated weighting rule. Both rules are evaluated and compared by instrumental and auditive tests. They succeed about equally well in attenuating the noise and the residual echo. In listening tests, however, the psychoacoustically motivated weighting rule is mostly preferred since it leads to more natural near end speech and to less annoying residual noise. Stefan Gustafsson, Rainer Martin 0001, Peter Jax, Peter Vary |
IEEE Trans. Speech Audio Process. | 2 |
| 2001 | Estimation of missing LSF parameters using Gaussian mixture modelsabstractSpeech transmission over packet networks has to cope with packet delays and packet losses. When a packet loss occurs the missing information must be estimated. We focus on restoring the spectral parameters of a speech coder. A novel approach to estimating missing line spectral frequency (LSF) parameters using Gaussian mixture models (GMM) is proposed. We present the estimation algorithm and study its performance when one or several LSF parameters are lost. We show that a GMM of a relatively low order is sufficient to achieve a substantial improvement in the parameter SNR. Therefore, the new estimation procedure requires much less memory than histogram based estimation methods. Rainer Martin 0001, Carsten Hoelper, Ingo Wittke |
ICASSP | 1 |
| 2001 | Planar superdirective microphone arrays for speech acquisition in the carabstractIn this paper we investigate a small broadside planar (2D) superdirective microphone array for speech acquisition in the car and compare its performance to linear arrays. The objective of this investigation is to replace an expensive directional microphone by a small array of inexpensive omnidirectional sensors. Since the array was designed to be used in the car environment it has to satisfy restrictions with respect to size and to the number of microphones. For all array configurations we present theoretical gains, actual measured gains using low-cost microphones, and beam patterns. For a fixed number of microphones and fixed array dimensions we show that the planar design leads to slightly superior array gains. Rainer Martin 0001, Alexey Petrovsky, Thomas Lotter |
INTERSPEECH | 1 |
| 2001 | Noise power spectral density estimation based on optimal smoothing and minimum statisticsabstractWe describe a method to estimate the power spectral density of nonstationary noise when a noisy speech signal is given. The method can be combined with any speech enhancement algorithm which requires a noise power spectral density estimate. In contrast to other methods, our approach does not use a voice activity detector. Instead it tracks spectral minima in each frequency band without any distinction between speech activity and speech pause. By minimizing a conditional mean square estimation error criterion in each time step we derive the optimal smoothing parameter for recursive smoothing of the power spectral density of the noisy speech signal. Based on the optimally smoothed power spectral density estimate and the analysis of the statistics of spectral minima an unbiased noise estimator is developed. The estimator is well suited for real time implementations. Furthermore, to improve the performance in nonstationary noise we introduce a method to speed up the tracking of the spectral minima. Finally, we evaluate the proposed method in the context of speech enhancement and low bit rate speech coding with various noise types. Rainer Martin 0001 |
IEEE Trans. Speech Audio Process. | 1 |
| 2000 | Optimized estimation of spectral parameters for the coding of noisy speechabstractIn this contribution we optimize a speech enhancement preprocessor such that a distortion measure in the line spectral frequency (LSF) domain is minimized. We can thus improve the estimation of spectral parameters of a speech coder when the input signal to the coder is a noisy speech signal. The optimization aims at the maximum noise reduction of the enhancement preprocessor. The average maximum noise reduction characteristic is determined as a function of the speech signal SNR and is approximated by an exponential function. Since LSF parameters are widely used in speech coding the results are applicable to a wide range of speech coders and enhancement preprocessors. We report experimental results for an MMSE log spectral amplitude estimator in conjunction with the new ETSI adaptive multi-rate (AMR) speech coder. We found that the method is most effective for the low bit rate coding modes. Rainer Martin 0001, Ingo Wittke, Peter Jax |
ICASSP | 1 |
| 1999 | Low delay analysis/synthesis schemes for joint speech enhancement and low bit rate speech codingabstractA corpus of spontaneous route descriptions was collected from 8 speakers (5 males and 3 females). The corpus was labelled according to the ToBI standard and a discourse analysis was completed. Four discourse tour acts were identified and these were found to occur mainly in non-embedded linear sequences. In general, the intonation of each route description was characterised by a single intonational phrase containing many intermediate phrases. There was a tendency for boundaries between tour acts and intermediate phrases to coincide, but there is usually more than one intermediate phrase to one tour act. Rainer Martin 0001, Hong-Goo Kang, Richard V. Cox |
EUROSPEECH | 1 |
| 1998 | Combined acoustic echo control and noise reduction for hands-free telephony
Stefan Gustafsson, Rainer Martin 0001, Peter Vary |
Signal Process. | 2 |
| 1997 | Combined acoustic echo control and noise reduction for mobile communications
Stefan Gustafsson, Rainer Martin 0001 |
EUROSPEECH | 2 |
| 1996 | The echo shaping approach to acoustic echo control
Rainer Martin 0001, Stefan Gustafsson |
Speech Commun. | 1 |
| 1995 | Coupled adaptive filters for acoustic echo control and noise reductionabstractPresents new adaptive algorithms for acoustic echo control and noise reduction which employ one, two, or possibly more microphone signals. The new algorithms accommodate high echo attenuation and lead to implementations with reduced complexity. These algorithms combine a conventional FIR echo canceller with a second NLMS-adapted FIR filter which attenuates residual echoes. The paper presents a one-microphone system with improved echo attenuation and a two-microphone system which attenuates acoustic echoes as well as ambient noise and near end speech reverberation. The algorithms can be interpreted as a frequency selective generalization of the well known voice controlled switch. The paper explains the algorithms and presents experimental results in real acoustic environments. Rainer Martin 0001, Jan Altenhöner |
ICASSP | 1 |
| 1995 | Design and optimization of a two microphone speech enhancement system
Rainer Martin 0001 |
EUROSPEECH | 1 |
| 1993 | An efficient algorithm to estimate the instantaneous SNR of speech signalsabstractThis contribution presents an efficient algorithm to estimate the instantaneous signal-to-noise ratio of speech signals. The algorithm is capable to track non stationary noise signals and has a low computational complexity. It does not need a speech activity detector nor histograms to learn signal statistics. The algorithm is based on the observation that a noise power estimate can be obtained using minimum values of a smoothed power estimate. This paper will present this algorithm, its performance, its limits, and some applications. Rainer Martin 0001 |
EUROSPEECH | 1 |