EDBT 2026 Demo / reviewers in the wild / expert
Sven Nordholm
dblp:09/497 · also Sven E. Nordholm, Sven Erik Nordholm
· DBLP profile ↗
98ranked-venue papers
5as first author
5since 2021 · last 2023
0000-0001-8942-5328ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 3 first-authorArtificial intelligence and machine learning · 31 · 2 first-author · 5 since 2021Computer networks · 14Security and privacy · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
19 papers |
Audio and music processing · 99% Image and video processing · 1% | |
| Computer networks
3 papers |
Wireless sensing and localization · 73% Physical-layer communications · 27% | |
| Artificial intelligence
2 papers |
Video understanding and tracking · 55% Probabilistic and Bayesian machine learning · 27% Speech recognition and synthesis · 18% | |
| Theoretical computer science
2 papers |
Coding theory · 62% Mathematical optimization · 38% |
Topics — the 30 heaviest of 44, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing › source separation
blind source separation |
1.2 | 4 | 2022 | Audio-Visual Based Online Multi-Source Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Blind Separation for Multiple Moving Sources With Labeled Random Finite Sets · IEEE ACM Trans. Audio Speech Lang. Process. 2021 Convolutive blind signal separation with post-processing · IEEE Trans. Speech Audio Process. 2004 |
Audio and music processing › sound source localization
acoustic source tracking |
1.1 | 2 | 2022 | Audio-Visual Based Online Multi-Source Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Blind Separation for Multiple Moving Sources With Labeled Random Finite Sets · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing › hearing aids
acoustic feedback cancellation |
0.8 | 2 | 2020 | Acoustic Feedback Suppression for Multi-Microphone Hearing Devices Using a Soft-Constrained Null-Steering Beamformer · IEEE ACM Trans. Audio Speech Lang. Process. 2020 Null-Steering Beamformer-Based Feedback Cancellation for Multi-Microphone Hearing Aids With Incoming Signal Preservation · IEEE ACM Trans. Audio Speech Lang. Process. 2019 |
Audio and music processing
hearing aids |
0.7 | 4 | 2022 | Two-Microphone Hearing Aids Using Prediction Error Method for Adaptive Feedback Control · IEEE ACM Trans. Audio Speech Lang. Process. 2018 A Class of Pareto Optimal Binaural Beamformers · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Acoustic Feedback Suppression for Multi-Microphone Hearing Devices Using a Soft-Constrained Null-Steering Beamformer · IEEE ACM Trans. Audio Speech Lang. Process. 2020 |
Audio and music processing
microphone array processing |
0.7 | 5 | 2022 | A Class of Pareto Optimal Binaural Beamformers · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Near-field broadband beamformer design via multidimensional semi-infinite-linear programming techniques · IEEE Trans. Speech Audio Process. 2003 Performance limits in subband beamforming · IEEE Trans. Speech Audio Process. 2003 |
Wireless sensing and localization
acoustic source localization |
0.7 | 1 | 2023 | Distributed Microphone Array Localization Problem via SDP-SOCP Method · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Wireless sensing and localization › acoustic source localization
microphone array localization |
0.7 | 1 | 2023 | Distributed Microphone Array Localization Problem via SDP-SOCP Method · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Wireless sensing and localization › range-based localization › time-based localization
time-difference-of-arrival localization |
0.7 | 1 | 2023 | Distributed Microphone Array Localization Problem via SDP-SOCP Method · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › bayesian inference
bayesian filtering |
0.6 | 1 | 2022 | A Bayesian Filter for Multi-View 3D Multi-Object Tracking With Occlusion Handling · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Video understanding and tracking
multi-object tracking |
0.6 | 1 | 2022 | A Bayesian Filter for Multi-View 3D Multi-Object Tracking With Occlusion Handling · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Computer vision › Video understanding and tracking › object tracking
occlusion handling |
0.6 | 1 | 2022 | A Bayesian Filter for Multi-View 3D Multi-Object Tracking With Occlusion Handling · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Audio and music processing › source separation
audio-visual sound separation |
0.6 | 1 | 2022 | Audio-Visual Based Online Multi-Source Separation · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Audio and music processing › speech enhancement
binaural beamforming |
0.6 | 1 | 2022 | A Class of Pareto Optimal Binaural Beamformers · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Audio and music processing › source separation
moving source separation |
0.5 | 1 | 2021 | Blind Separation for Multiple Moving Sources With Labeled Random Finite Sets · IEEE ACM Trans. Audio Speech Lang. Process. 2021 |
Audio and music processing
speech enhancement |
0.5 | 9 | 2020 | Acoustic Feedback Suppression for Multi-Microphone Hearing Devices Using a Soft-Constrained Null-Steering Beamformer · IEEE ACM Trans. Audio Speech Lang. Process. 2020 Null-Steering Beamformer-Based Feedback Cancellation for Multi-Microphone Hearing Aids With Incoming Signal Preservation · IEEE ACM Trans. Audio Speech Lang. Process. 2019 Noise Statistics Update Adaptive Beamformer With PSD Estimation for Speech Extraction in Noisy Environment · IEEE Trans. Speech Audio Process. 2008 |
Audio and music processing
beamforming |
0.4 | 2 | 2014 | On the Indoor Beamformer Design With Reverberation · IEEE ACM Trans. Audio Speech Lang. Process. 2014 Design of Steerable Spherical Broadband Beamformers With Flexible Sensor Configurations · IEEE Trans. Speech Audio Process. 2013 |
Mathematical optimization
convex relaxation |
0.2 | 1 | 2023 | Distributed Microphone Array Localization Problem via SDP-SOCP Method · IEEE ACM Trans. Audio Speech Lang. Process. 2023 |
Natural language and speech › Speech recognition and synthesis › speech enhancement › beamforming
multichannel wiener filter |
0.2 | 1 | 2014 | Effective binaural multi-channel processing algorithm for improved environmental presence · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Natural language and speech › Speech recognition and synthesis
speech enhancement |
0.2 | 1 | 2014 | Effective binaural multi-channel processing algorithm for improved environmental presence · IEEE ACM Trans. Audio Speech Lang. Process. 2014 |
Audio and music processing › spatial audio
binaural cue preservation |
0.2 | 1 | 2022 | A Class of Pareto Optimal Binaural Beamformers · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Audio and music processing
acoustic echo cancellation |
0.2 | 2 | 2012 | A Spectral Slit Approach to Doubletalk Detection · IEEE Trans. Speech Audio Process. 2012 Adaptive microphone array employing calibration signals: an analytical evaluation · IEEE Trans. Speech Audio Process. 1999 |
Audio and music processing › beamforming
spherical array beamforming |
0.2 | 1 | 2013 | Design of Steerable Spherical Broadband Beamformers With Flexible Sensor Configurations · IEEE Trans. Speech Audio Process. 2013 |
Physical-layer communications
equalization |
0.2 | 1 | 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM Systems · IEEE Trans. Commun. 2013 |
Physical-layer communications › modulation › multicarrier modulation
OFDM |
0.2 | 1 | 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM Systems · IEEE Trans. Commun. 2013 |
Physical-layer communications › equalization
turbo equalization |
0.2 | 1 | 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM Systems · IEEE Trans. Commun. 2013 |
Coding theory › error-correcting codes › decoding › iterative decoding
factor graphs |
0.2 | 1 | 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM Systems · IEEE Trans. Commun. 2013 |
Coding theory › error-correcting codes › decoding
iterative decoding |
0.2 | 1 | 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM Systems · IEEE Trans. Commun. 2013 |
Audio and music processing › acoustic echo cancellation
double-talk detection |
0.1 | 1 | 2012 | A Spectral Slit Approach to Doubletalk Detection · IEEE Trans. Speech Audio Process. 2012 |
Physical-layer communications › signal detection
MIMO detection |
0.1 | 1 | 2012 | A Fourier Based Method for Approximating the Joint Detection Probability in MIMO Communications · IEEE Trans. Commun. 2012 |
Audio and music processing › speech recognition
robust speech recognition |
0.1 | 1 | 2011 | A New Evidence Model for Missing Data Speech Recognition With Applications in Reverberant Multi-Source Environments · IEEE Trans. Speech Audio Process. 2011 |
Methods — techniques the papers use, named apart from their topics
semidefinite programming · 1.3second-order cone programming · 1.3relaxation · 1.3labeled random finite set · 1.1generalized side-lobe canceller · 1.1min-max optimization · 0.8least-squares optimization · 0.8multi-objective optimization · 0.6mean squared error cost function · 0.6bayesian multi-view multi-object filtering · 0.63d occlusion model · 0.6steered-response power phase transform · 0.5soft-input soft-output equalization · 0.3relative transfer function · 0.3prediction error method · 0.3forney-style factor graph · 0.3adaptive filter · 0.3speech presence probability · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Distributed Microphone Array Localization Problem via SDP-SOCP MethodabstractIn multimedia applications, it is common to employ acoustic sensors collectively to enhance signals and to locate sound sources. A direct problem can be formulated to locate sound sources from a set of known sensors. In order to form the acoustic sensor network, it is important to locate the sensor array locations first. However, unlike other networks in which direct time-of-arrival (TOA) measurements might be possible, acoustic distributed network can only obtain time-difference-of-arrival (TDOA) measures indirectly from various sound source anchors. While it is common to employ convex optimization techniques to localize sensor locations in a network with TOA information, it has not been studied properly when it comes to TDOAs. This paper considers the microphone array localization problem in a distributed acoustic network with TDOA measurements. We formulate the inverse problem which applied the known source locations to identify the wireless array configuration and estimate the location for each array. The proposed method formulates a mixed semidefinite programming (SDP) and second-order cone programming (SOCP) relaxation model, and then the acoustic geometry is obtained by solving a linear optimal programming. Furthermore, the characteristics of the optimal solution are studied and exact relaxation conditions are given. Experimental results demonstrate that the proposed mixed model can successfully estimate the sensor locations in noisy and reverberant environments for 2-dimensional and 3-dimensional space, which outperforms other relaxation methods. He Qi, Ka Fai Cedric Yiu, Sven Nordholm |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | A Bayesian Filter for Multi-View 3D Multi-Object Tracking With Occlusion HandlingabstractThis paper proposes an online multi-camera multi-object tracker that only requires monocular detector training, independent of the multi-camera configurations, allowing seamless extension/deletion of cameras without retraining effort. The proposed algorithm has a linear complexity in the total number of detections across the cameras, and hence scales gracefully with the number of cameras. It operates in the 3D world frame, and provides 3D trajectory estimates of the objects. The key innovation is a high fidelity yet tractable 3D occlusion model, amenable to optimal Bayesian multi-view multi-object filtering, which seamlessly integrates, into a single Bayesian recursion, the sub-tasks of track management, state estimation, clutter rejection, and occlusion/misdetection handling. The proposed algorithm is evaluated on the latest WILDTRACKS dataset, and demonstrated to work in very crowded scenes on a new dataset. Jonah Ong, Ba-Tuong Vo, Ba-Ngu Vo, Du Yong Kim, Sven Nordholm |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | A Class of Pareto Optimal Binaural BeamformersabstractThe objective of binaural multi-microphone speech enhancement algorithms can be viewed as a multi-criteria design problem as there are several requirements to be met. The objective is not only to extract the target speaker without distortion, but also to suppress interfering sources (e.g., competing speakers) and ambient background noise, while preserving the auditory impression of the complete acoustic scene. Such a multi-objective problem (MOP) can be solved using a Pareto frontier, which provides a useful trade-off between the different criteria. In this paper, we propose a unified Pareto optimization framework, which is achieved by defining a generalized mean squared error (MSE) cost function, derived from a MOP. The solution to the multi-criteria problem is grounded on a solid mathematical foundation. The MSE cost function consists of a weighted sum of speech distortion (SD), partial interference reduction (IR), and partial noise reduction (NR) terms with scaling parameters that control the amount of IR and NR. The filter minimizing this generalized cost function, denoted Pareto optimal binaural multichannel Wiener filter (Pareto-BMWF), constitutes a generalization of various binaural MWF-based and binaural MVDR-based beamformers. This solution is optimal for any set of parameters. The improved speech enhancement capabilities are experimentally demonstrated using real-signal recordings when estimation errors are present and the binaural cue preservation capabilities are analyzed. Elior Hadad, Simon Doclo, Sven Nordholm, Sharon Gannot |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Audio-Visual Based Online Multi-Source SeparationabstractMeeting or conference assistance is a popular application that typically requires compact configurations of co-located audio and visual sensors. This paper proposes a novel solution for online separation of an unknown and time-varying number of moving sources using only a single microphone array co-located with a single visual device. The approach exploits the complementary nature of simultaneous audio and visual measurements, accomplished by a model-centric 3-stage process of detection, tracking, and (spatial) filtering, which performs separation in a block-wise or recursive fashion. Fusing the measurements requires solving the multi-modal space-time permutation problem, since the audio and visual measurements reside in different observation spaces, but also are unidentified or unlabeled (with respect to the unknown and time-varying number of sources), and are subject to noise, extraneous measurements and missing measurements. A labeled random finite set tracking filter is applied to resolve the permutation problem and recursively estimate the source identities and trajectories. A time-varying set of generalized side-lobe cancellers is constructed based on the tracking estimates to perform online separation. Evaluations are undertaken with live human speakers. Jonah Ong, Ba-Tuong Vo, Sven Nordholm, Ba-Ngu Vo, Diluka Moratuwage, Changbeom Shim |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Blind Separation for Multiple Moving Sources With Labeled Random Finite SetsabstractThis paper proposes a novel solution for separating an unknown and time-varying number of moving acoustic sources in a blind setting using multiple microphone arrays. A standard steered-response power phase transform method is applied to extract source position measurements, which inevitably contain noise, false detections, missed detections, and are not labeled with the source identities. The imperfect measurements lead to the space-time permutation problem, as there is no information on how the measurements are associated to the sources in space, nor how the measurements are connected across time, if at all. To solve this problem, a labeled random finite set tracking framework is adopted to jointly estimate the source positions and their labels or identities. Based on these trajectory estimates, a corresponding set of time-varying generalized side-lobe cancellers is constructed to perform source separation. The overall algorithm operates in a block-wise or an online fashion and is scalable with the number of microphone arrays. The quality of the measurements, tracking, and separation, are evaluated respectively, with the OSPA metric, OSPA(2)metric, and ITU-T P.835 based listening tests, on both real-world and simulated data. Jonah Ong, Ba-Tuong Vo, Sven Nordholm |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Acoustic Feedback Suppression for Multi-Microphone Hearing Devices Using a Soft-Constrained Null-Steering BeamformerabstractAcoustic feedback occurs in hearing aids due to the coupling between the hearing aid loudspeaker and microphone(s). In order to reduce the acoustic feedback, adaptive filters are commonly used to estimate the feedback contribution in the microphone(s). While theoretically allowing for perfect feedback cancellation, in practice the adaptive filter typically converges to a biased optimal solution due to the closed-loop acoustical system of the hearing aid. Previously it has therefore been proposed to suppress the acoustic feedback contribution for an earpiece with multiple integrated microphones and loudspeakers using a fixed null-steering beamformer and hence avoiding a biased adaption. While previous null-steering beamforming approaches aimed at perfect preservation of the incoming signal using its relative transfer function (RTF), in this article we propose to use a soft constraint that allows to trade off between incoming signal preservation and feedback suppression. We formulate the computation of the beamformer coefficients both as a least-squares optimization procedure, aiming to minimize the residual feedback power, and as a min-max optimization procedure, aiming to directly maximize the maximum stable gain of the hearing aid. Experimental evaluations were performed using measured acoustic feedback paths from a custom earpiece with two microphones in the vent and a third microphone in the concha. Results show that the proposed fixed null-steering beamformer using the RTF-based soft constraint provides a reduction of the acoustic feedback by 7-8 dB compared to the previously proposed RTF-based hard constraint while limiting the distortions of the incoming signal in the beamformer output. Henning F. Schepker, Sven Nordholm, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2019 | Cross Evaluation of Speech Enhancement Methods under Different Noise ConditionsabstractIn this paper, we present a cross evaluation of different a priori SNR estimation methods as well as different time-frequency analysis processing using a subjective listening test. The noisy signal is corrupted by different types of background noise i.e. babble and pink, and varying levels of input SNR (0 dB and 10 dB). The signals are processed using Short Time Fourier Transform (STFT) or Critical Band (CB) processing. After estimating the clean speech signal, it was presented to 10 participants for evaluation using a subjective listening test according to (ITU-TP.835) methodology. The results demonstrate that the participants preferred the speech signal processed using CB for low SNR levels and non-stationary background noise, which means that critical band based frequency scale is more useful in adverse noisy conditions. Lara Nahma, Pei Chee Yong, Hai Huyen Dam, Sven Nordholm |
ICASSP | 4 |
| 2019 | Null-Steering Beamformer-Based Feedback Cancellation for Multi-Microphone Hearing Aids With Incoming Signal PreservationabstractIn hearing aids, acoustic feedback occurs due to the coupling between the hearing aid loudspeaker and microphone(s). In order to reduce the acoustic feedback, adaptive filters are commonly used to estimate the feedback contribution in the microphone(s). While theoretically allowing for perfect feedback cancellation, in practice the adaptive filter converges to an optimal solution that is typically biased due to the closed-loop acoustical system of the hearing aid. In order to avoid the adaptation to a biased optimal solution, in this paper we propose to use a fixed beamformer to cancel the acoustic feedback contribution for an earpiece with multiple integrated microphones and loudspeakers. By steering a spatial null in the direction of the hearing aid loudspeaker, we show that theoretically perfect feedback cancellation can be achieved. While previous null-steering beamforming approaches did not control for distortions of the incoming signal, in this paper we propose to incorporate a constraint based on the relative transfer function (RTF) of the incoming signal, aiming to perfectly preserve this signal. We formulate the computation of the beamformer coefficients both as a least-squares optimization procedure, aiming to minimize the residual feedback power, and as a min-max optimization procedure, aiming to directly maximize the maximum stable gain of the hearing aid. Experimental results using measured acoustic feedback paths from a custom earpiece with two microphones in the vent and a third microphone in the concha show that the proposed fixed null-steering beamformer using the RTF-based constraint provides a reduction of the acoustic feedback and substantially increases the added stable gain while preserving the incoming signal. This can even be achieved for unknown acoustic feedback paths and incoming signal directions. Henning F. Schepker, Sven Nordholm, Linh Thi Thuc Tran, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2018 | Bayesian noise estimation in the modulation domain
Maneesh K. Singh, Siow Yong Low, Sven Nordholm, Zhuquan Zang |
Speech Commun. | 3 |
| 2018 | Two-Microphone Hearing Aids Using Prediction Error Method for Adaptive Feedback ControlabstractA challenge in hearing aids is adaptive feedback control which often uses an adaptive filter to estimate the feedback path. This estimate of the feedback path usually results in a bias due to the correlation between the loudspeaker signal and the incoming signal. The prediction error method (PEM) is a popular method for reducing this bias for adaptive feedback control (AFC) in hearing aids, providing a significant performance improvement compared to conventional adaptive feedback control techniques. However, the PEM-based AFC (PEM-AFC) applications are still limited to single-microphone single-loudspeaker (SMSL) systems. This paper investigates the application of the PEM-AFC to a two-microphone single-loudspeaker hearing aid with detailed theoretical analysis as well as practical experiments. In the proposed method, PEM-AFC2, we use the two-microphone adaptive feedback control (AFC2) method with two microphones and one loudspeaker. The incoming signals at the two microphones are related by a relative transfer function (RTF) which is used to predict the incoming signal at the main microphone. In addition, a prefilter is employed to prewhiten the loudspeaker and the microphone signals before the adaptive filter estimates. As a result, the proposed method obtains a lower bias and a faster tracking rate compared to the PEM-AFC and the AFC2 method, while still maintaining a good quality of the incoming signal. A new derivation for optimal filters in the AFC2 method will also be provided. The performance of the proposed method is evaluated for speech shaped noise as incoming signal and with undermodeling the RTF as well as with perfect modeling the RTF. Moreover, different types of incoming signals and a sudden change of feedback paths are also considered. The experimental results show that the proposed approach yields a significant performance improvement compared to existing state-of-the-art AFC methods such as the PEM-AFC and the AFC2. Linh Thi Thuc Tran, Sven Nordholm, Henning F. Schepker, Hai Huyen Dam, Simon Doclo |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2017 | A LabVIEW-based implementation of real-time underwater acoustic OFDM systemabstractThe orthogonal frequency-division multiplexing (OFDM) technology receives increasing attention in underwater acoustic (UA) communications. This paper presents a real-time OFDM-based UA communication system, implemented using the National Instruments CompactDAQ device and the LabVIEW software. The system design including both the transmitter and receiver is discussed. The performance of this real-time system is verified through a UA communication experiment conducted recently in a tank. Compared with conventional digital signal processor (DSP)-based design, the proposed implementation simplifies the prototype design and reduces the software development time. Peng Chen 0059, Yue Rong, Sven Nordholm, Alec J. Duncan, Zhiqiang He 0001 |
APCC | 3 |
| 2017 | Improved a priori snr estimation in speech enhancementabstractIn this paper, a modified a priori SNR estimation method with application in speech enhancement is presented. The decision directed (DD) approach for the a priori SNR estimation has been the most popular method due to its ability to eliminate musical noise in speech enhancement. However, a limitation of this method is the slow tracking of speech onsets that generates a transient distortion to the speech. The proposed method modifies the DD a priori SNR estimator by employing an adaptive smoothing factor based on the instan-taneous a posteriori SNR and an adaptive sigmoid function to obtain an improved tracking of speech onsets. Experimental results using several instrumental measures show the ability of the proposed method to preserve weak speech components when compared to two reference methods while maintaining the advantage of the DD approach in eliminating the musical noise. Lara Nahma, Pei Chee Yong, Hai Huyen Dam, Sven Nordholm |
APCC | 4 |
| 2017 | Convex combination framework for a priori SNR estimation in speech enhancementabstractThe paper proposes a convex combination fusion function based on a sigmoid function for the estimation of the a priori SNR in a speech enhancement framework with critical frequency band processing. The proposed method does not only eliminate the one frame delay generated by the well-known decision directed approach but also increases the adaptation speed during abrupt changes in the SNR estimation. As a result, the advantage of low musical noise has been maintained while more weak speech components have been preserved. Experimental results using instrumental and subjective measures also indicate improvement in speech quality compared to the reference methods. Lara Nahma, Pei Chee Yong, Hai Huyen Dam, Sven Nordholm |
ICASSP | 4 |
| 2017 | Null-steering beamformer for acoustic feedback cancellation in a multi-microphone earpiece optimizing the maximum stable gainabstractCommonly adaptive filters are used to reduce the acoustic feedback in hearing aids. While theoretically allowing for perfect cancellation of the feedback signal, in practice the adaptive filter solution is typically biased due to the closed-loop hearing aid system. In contrast to conventional behind-the-ear hearing aids, in this paper we consider an earpiece with multiple integrated microphones. For such an earpiece it has previously been proposed to use a fixed beamformer to reduce the acoustic feedback in the microphones which has been designed to minimize a least-squares cost function. In this paper we propose to design the beamformer by minimizing a min-max cost function which directly maximizes the maximum stable gain of the earpiece. Furthermore, we propose a robust extension of the min-max cost function maximizing the worst-case maximum stable gain over a set of acoustic feedback paths. Experimental results using measured acoustic feedback paths show that the feedback cancellation performance of the fixed beamformer can be considerably improved by minimizing the proposed min-max optimization problem, while maintaining a high perceptual quality of the incoming signal. Henning F. Schepker, Linh Thi Thuc Tran, Sven Nordholm, Simon Doclo |
ICASSP | 3 |
| 2017 | Proportionate NLMS for adaptive feedback control in hearing aidsabstractThe proportionate normalized least-mean-squares (PNLMS) algorithm is commonly used in acoustic echo cancellation (AEC) context. It provides faster initial convergence and tracking rates compared to the NLMS algorithm for the case of sparse echo impulse responses. The improved PNLMS algorithm (IPNLMS) has been proven to be more powerful than PNLMS by exploiting new rules for computing the weight of each step-size corresponding to each adaptive filter coefficient. However, the application of the PNLMS and the IPNLMS algorithms for adaptive feedback control (AFC) in hearing aids (HAs) is still limited due to high correlation between the loudspeaker and incoming signals. This paper proposes implementations of the PNLMS/IPNLMS algorithms for AFC using the prediction error method (PEM) for hearing aids. The proposed methods have been evaluated for both speech and music incoming signals. Simulation shows that the proposed methods have faster initial convergence and tracking than the PEM using the NLMS algorithm (PEM-NLMS). Linh Thi Thuc Tran, Henning F. Schepker, Simon Doclo, Hai Huyen Dam, Sven Nordholm |
ICASSP | 5 |
| 2017 | Joint Channel Estimation and Impulsive Noise Mitigation in Underwater Acoustic OFDM Communication SystemsabstractImpulsive noise occurs frequently in underwater acoustic (UA) channels and can significantly degrade the performance of UA orthogonal frequency-division multiplexing (OFDM) systems. In this paper, we propose two novel compressed sensing based algorithms for joint channel estimation and impulsive noise mitigation in UA OFDM systems. The first algorithm jointly estimates the channel impulse response and the impulsive noise by utilizing pilot subcarriers. The estimated impulsive noise is then converted to the time domain and removed from the received signals. We show that this algorithm reduces the system bit-error-rate through improved channel estimation and impulsive noise mitigation. In the second proposed algorithm, a joint estimation of the channel impulse response and the impulsive noise is performed by exploiting the initially detected data. Then, the estimated impulsive noise is removed from the received signals. The proposed algorithms are evaluated and compared with existing methods through numerical simulations and on real data collected during a UA communication experiment conducted in the estuary of the Swan River, WA, Australia, during December 2015. The results show that the proposed approaches consistently improve the accuracy of channel estimation and the performance of impulsive noise mitigation in UA OFDM communication systems. Peng Chen 0059, Yue Rong, Sven Nordholm, Zhiqiang He 0001, Alexander J. Duncan |
IEEE Trans. Wirel. Commun. | 3 |
| 2016 | Improving adaptive feedback cancellation in hearing aids using an affine combination of filtersabstractIn adaptive feedback cancellation an adaptive filter is used to model the acoustic feedback path between the hearing aid loudspeaker and the microphone. An important parameter for adaptive filters is the step-size, providing a trade-off between fast convergence and low steady-state misalignment. In order to achieve both fast convergence as well as low steady-state misalignment, it has been proposed to use an affine combination scheme of two filters operating with different step-sizes. In this paper we apply such an affine combination scheme to the acoustic feedback cancellation problem in hearing aids. We show that for speech signals a time-domain affine combination scheme yields a biased solution. To reduce this bias we propose to use a partitioned-block frequency-domain affine combination scheme. Experimental results using measured acoustic feedback paths show that in terms of misalignment and added stable gain the proposed adaptive feedback cancellation system outperforms a system that only uses a single adaptive filter with either of the fixed step-sizes used for the affine combination scheme. Henning F. Schepker, Linh Thi Thuc Tran, Sven Nordholm, Simon Doclo |
ICASSP | 3 |
| 2016 | Simplified MMSE Precoding Design in Interference Two-Way MIMO Relay SystemsabstractWe investigate the transceiver design for interference two-way amplify-and-forward multiple-input multiple-output relay communication systems. A novel algorithm with a closed-form solution is developed to optimize the relay precoding matrix based on its optimal structure and a modified transmission power constraint at the relay node. An iterative algorithm is proposed to minimize the sum mean-squared error of the signal waveform estimation. Simulation results demonstrate that the proposed algorithm achieves a better performance-complexity tradeoff compared with existing techniques. Khoa Xuan Nguyen, Yue Rong, Sven Nordholm |
IEEE Signal Process. Lett. | 3 |
| 2016 | On a New SDP-SOCP Method for Acoustic Source Localization ProblemabstractAcoustic source localization has many important applications. Convex relaxation provides a viable approach of obtaining good estimates very efficiently. There are two popular convex relaxation methods using either semi-definite programming (SDP) or second-order cone programming (SOCP). However, the performances of the methods have not been studied properly in the literature and there is no comparison in terms of accuracy and performance. The aims of this article are twofold. First of all, we study and compare several convex relaxation methods. We demonstrate, by numerical examples, that most of the convex relaxation methods cannot localize the source exactly, even in the performance limit when the time difference of arrival (TDOA) information is exact. In addressing this problem, we propose a novel mixed SDP-SOCP relaxation model and study the characteristics of the optimal solutions and its localizable region. Furthermore, an error correction scheme for the proposed SDP-SOCP model is developed so that exact localization can be achieved in the performance limit. Experimental data have been collected in a room with two different array configurations to demonstrate our proposed approach. Ka Fai Cedric Yiu, Sven Nordholm, Yinyu Ye 0001 |
ACM Trans. Sens. Networks | 3 |
| 2015 | A low complexity iterative soft-decision feedback MMSE-PIC detection algorithm for massive MIMOabstractIn MIMO applications, the minimum mean square error parallel interference cancellation (MMSE-PIC) based Soft-Input Soft-Output (SISO) detector has been widely adopted because of its low complexity and good bit error rate (BER) performance. In this paper, we firstly propose to use a Gaussian model based MMSE detection algorithm to implement MMSE-PIC with low complexity. This algorithm, which can detect a length-Nrreceived data block by a single Hermitian matrix (sized Nt× Nt) inversion, is especially preferable in Massive MIMO up-link applications where the number of transmit antennas Ntfrom each end terminal is much less than the number of receive antennas Nrin the Base Station. Then we derive a new method to calculate the matrix inversion by a linear combination of two matrices, which reduces the complexity from O(Nt3) to O(Nt2). At last, in order to improve the system performance for the first pass when there is no a priori information available, a self-iteration method is proposed and thus a system performance gain of 1dB to 2dB is achieved at the cost of modest complexity increase. Licai Fang, Lu Xu 0003, Qinghua Guo 0001, Defeng Huang, Sven Nordholm |
ICASSP | 5 |
| 2015 | Assistive listening headsets for high noise environments: Protection and communicationabstractIn industrial noise environments, the use of assistive listening headsets is a means to provide adequate access to voice communication while wearing hearing protection. This paper presents a performance evaluation and comparison of two different methods to provide the binaural speech enhancement in real industrial noise scenarios. The investigated binaural methods based on differential beamforming and multichannel Wiener filter show different strengths and weaknesses. A transient noise suppression algorithm is also proposed and evaluated. Performance evaluation shows that this algorithm, together with the binaural multi-channel Wiener filter approach, can successfully reduce the hammering noise. This can be observed from the PESQ scores and the signal characteristics. Sven Nordholm, Alan Davis, Pei Chee Yong, Hai Huyen Dam |
ICASSP | 1 |
| 2015 | Analysis of Two Microphone Method for Feedback CancellationabstractAcoustic feedback cancellation in hearing aids makes use of adaptive filters to continuously identify and track variations to the feedback path. One of the biggest problems remaining in using adaptive filters for feedback cancellation is the biased estimation of the filter's coefficients. In order to remove the undesired correlation between the loudspeaker and incoming signal, a recent alternative scheme proposed to employ an additional microphone. This microphone can provide added information to obtain an incoming signal estimate. This estimate is removed from the primary microphone signal to create the error signal which adapts the canceler's coefficients. This letter provides the theoretical analysis for the two microphone method. It presents analytic expressions showing that the optimal solution is no longer dependent on the signal correlation aforementioned but is now mainly determined by the additional feedback path. Finally, it demonstrates simulation results with the prediction error method in terms of misalignment and maximum gain for a proposed microphone placement. The results show that a more stable solution is obtained with the proposed two microphone approach. Carlos Renato C. Nakagawa, Sven Nordholm, Wei-Yong Yan |
IEEE Signal Process. Lett. | 2 |
| 2015 | MMSE-Based Transceiver Design Algorithms for Interference MIMO Relay SystemsabstractIn this paper, we investigate the transceiver design for amplify-and-forward interference multiple-input multiple-output (MIMO) relay communication systems, where multiple transmitter-receiver pairs communicate simultaneously with the aid of a relay node. The aim is to minimize the mean-squared error (MSE) of the signal waveform estimation at the receivers subjecting to transmission power constraints at the transmitters and the relay node. As the transceiver optimization problem is nonconvex with matrix variables, the globally optimal solution is intractable to obtain. To overcome the challenge, we propose an iterative transceiver design algorithm where the transmitter, relay, and receiver matrices are optimized iteratively by exploiting the optimal structure of the relay precoding matrix. To reduce the computational complexity of optimizing the relay precoding matrix, we propose a simplified relay matrix design through modifying the transmission power constraint at the relay node. The modified relay optimization problem has a closed-form solution. Simulation results demonstrate that the proposed algorithms perform better than the existing techniques in terms of both MSE and bit-error-rate. Khoa Xuan Nguyen, Yue Rong, Sven Nordholm |
IEEE Trans. Wirel. Commun. | 3 |
| 2014 | On the use of contextual time-frequency information for full-band clustering-based convolutive blind source separationabstractIn this paper we propose to incorporate contextual time-frequency information for clustering-based blind source separation. Previous clustering-based approaches have successfully used clustering techniques to estimate time-frequency separation masks; however, these approaches generally do not consider the contextual information of each time-frequency slot. Motivated by the homogenous behavior of speech signals, we modify the fuzzy c-means clustering to bias the results in favor of cluster membership homogeneity within localized neighborhoods in the time-frequency space. Experimental evaluations in both simulated and real-world underdetermined environments demonstrate improvement in source separation performance over previous clustering approaches. Matt Atcheson, Ingrid Jafari, Roberto Togneri, Sven Nordholm |
ICASSP | 4 |
| 2014 | Source number estimation in reverberant conditions via full-band weighted, adaptive fuzzy c-means clusteringabstractWe introduce a novel approach for source number estimation through an adaptive fuzzy c-means clustering. Spatial feature vectors are extracted from microphone observations, weighted for reliability and then clustered in a full-band manner using an adaptive variation on the fuzzy c-means. A number of quality measures are combined to produce a weighted sum which is used to find the optimal number of clusters at each iteration of the clustering algorithm. Experimental evaluations using real-world recordings from a reverberant room (RT60= 390 ms) demonstrated encouraging performance in both even- and under-determined conditions. Joshua Hollick, Ingrid Jafari, Roberto Togneri, Sven Nordholm |
ICASSP | 4 |
| 2014 | Closed-loop feedback cancellation utilizing two microphones and transform domain processingabstractIn this paper we are studying the use of two microphones for acoustic feedback cancellation in hearing aids. With the two microphones approach, an additional microphone is employed to provide added information about the signals which is then utilized to obtain an incoming signal estimate. This estimate is removed from the error signal prior to adapting the canceler, thus removing the undesired signal correlation. In this paper, we propose to use orthogonal transforms with the two microphones approach. The discrete Fourier transform and the discrete cosine transform are implemented to transform the adaptive filter signals. Also, a bank of adaptive filters is employed, each adapting to different portions of the spectrum for a finer control of the adaptation process. Simulation results based on real measured feedback paths and speech signals show improved convergence rates and stable solutions. Carlos Renato C. Nakagawa, Sven Nordholm, Felix Albu, Wei-Yong Yan |
ICASSP | 2 |
| 2014 | Exploiting cyclic prefix for joint detection, decoding and channel estimation in OFDM via EM algorithm and message passingabstractThis paper considers the coded OFDM system and instead of discarding the cyclic prefix (CP) at the receiver, we utilize the CP observation for joint detection, decoding and channel estimation. In particular, detection and decoding are performed iteratively between an equalizer and a soft-input soft-output (SISO) decoder based on the turbo principle, and the expectation-maximization (EM) algorithm is employed in the equalizer for joint detection and channel estimation via message passing. Models for the CP observation, non-CP observation and the time correlation of the time-varying channel are presented in Forney-style factor graphs (FFGs), and a scheduling scheme is proposed to pass messages between the graphs. Simulation results show that with unknown channel impulse response (CIR), the performance of the proposed algorithm approaches the case where CIR is perfectly known and through proper exploitation of the CP, the proposed algorithm outperforms the conventional algorithm (i.e. CP is discarded) with known CIR, as well as the alternative algorithm in the literature (where CP is exploited) with unknown CIR. Jindan Yang, Qinghua Guo 0001, Defeng Huang, Sven Nordholm |
ICC | 4 |
| 2014 | On the use of the Watson mixture model for clustering-based under-determined blind source separationabstractCopyright © 2014 ISCA. In this paper, we investigate the application of a generative clustering technique for the estimation of time-frequency source separation masks. Recent advances in time-frequency clustering-based approaches to blind source separation have touched upon the Watson mixture model (WMM) as a tool for source separation. However, most methods have been frequency bin-wise and have thus required the additional permutation alignment stage, and previous full-band methods which employ the WMM have yet to be applied to the under-determined setting. We propose to evaluate the clustering ability of the WMM within the clustering-based source separation framework. Evaluations confirm the superiority of the WMM against other previously used clustering techniques such as the fuzzy c-means. Ingrid Jafari, Roberto Togneri, Sven Nordholm |
INTERSPEECH | 3 |
| 2014 | A Hybrid Descent Method for Optimal Sigmoid Filter DesignabstractIn this letter, a hybrid descent method is used to determine a set of filter parameters for a sigmoid filter which attempts to work under various SNR conditions. It overcomes the limitations of the current sigmoid filters that performs effectively only at a single SNR. Results show that significant improvement in terms of better speech qualities can be achieved by the proposed sigmoid filter when working under various SNR conditions. Kit Yan Chan, Sven Nordholm, Siow Yong Low, Pei Chee Yong, Ka Fai Cedric Yiu |
IEEE Signal Process. Lett. | 2 |
| 2014 | Feedback Cancellation With Probe Shaping CompensationabstractAdaptive feedback cancellation methods may integrate the use of probe signals to assist with the biased optimal solution in acoustic systems working in closed-loop. However, injecting a probe noise in the loudspeaker decreases the signal quality perceived by users of assistive listening devices. To counter this, probe signals are usually shaped to provide some level of perceptual masking. In this letter we show the impact of using a shaping filter on the system behavior in terms of convergence rate and steady state error. From this study, it can be concluded that shaping the probe signal may have detrimental influence in terms of system performance. Accordingly, we propose to use the unshaped probe signal combined with an inverse filter of the shaping filter to identify the feedback channel. This restructure of the problem restores convergence rate of LMS type algorithms. Furthermore, we also show that an adequate forward path delay is required to obtain an unbiased solution and that the suggested scheme reduces this delay. Carlos Renato C. Nakagawa, Sven Nordholm, Wei-Yong Yan |
IEEE Signal Process. Lett. | 2 |
| 2014 | On the Indoor Beamformer Design With ReverberationabstractBeamforming remains to be an important technique for signal enhancement. For applications in open space, the transfer function describing waves propagation has an explicit expression, which can be employed for beamformer design. However, the function becomes very complex in an indoor environment due to the effects of reverberation. In this paper, this problem is discussed. A method based on the image source method (ISM) is applied to model the room impulse responses (RIRs), which will act as the transfer function between source and sensor. The indoor beamformer design problem is formulated as a minimax optimization problem. We propose and study several optimization models based on the L1-norm to design the beamformer. We found that it is advantageous to separate early and late reverberations in the design process and better designs can be achieved. Several numerical experiments are presented using both simulated data and real recordings to evaluate the proposed methods. Zhibao Li, Ka Fai Cedric Yiu, Sven Nordholm |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2014 | Effective binaural multi-channel processing algorithm for improved environmental presenceabstractBinaural noise-reduction algorithms based on multi-channel Wiener filter (MWF) are promising techniques to be used in binaural assistive listening devices. The real-time implementation of the existing binaural MWF methods, however, involves challenges to increase the amount of noise reduction without imposing speech distortion, and at the same time preserving the binaural cues of both speech and noise components. Although significant efforts have been made in the literature, most developed methods so far have focused only on either the former or latter problem. This paper proposes an alternative binaural MWF algorithm that incorporates the non-stationarity of the signal components into the framework. The main objective is to design an algorithm that would be able to select the sources that are present in the environment. To achieve this, a modified speech presence probability (SPP) and a single-channel speech enhancement algorithm are utilized in the formulation. The resulting optimal filter also avoids the poor estimation of the second-order clean speech statistics, which is normally done by simple subtraction. Theoretical analysis and performance evaluation using realistic recorded data shows the advantage of the proposed method over the reference MWF solution in terms of the binaural cues preservation, as well as the noise reduction and speech distortion. Pei Chee Yong, Sven Nordholm, Hai Huyen Dam |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2013 | Incorporating multi-channel Wiener filter with single-channel speech enhancement algorithmabstractThe real-time implementation of the existing multi-channel Wiener filter (MWF) algorithms suffer from performance degradation due to the lack of robustness against estimation errors of the second-order statistics. The reasons are twofold: one, the estimation of the statistics relies on real voice activity detector (VAD), which often fails in adverse environments. Second, the MWF solutions involve estimation of the second order clean speech statistics, which also exaggerates the errors. This paper presents an MWF algorithm that requires neither VAD nor clean speech statistics. Performance evaluation under real scenarios shows that the proposed method outperforms the conventional MWF solution in terms of the trade-off between noise reduction and speech distortion. Pei Chee Yong, Sven Nordholm, Hai Huyen Dam, Yee Hong Leung, Chiong-Ching Lai |
ICASSP | 2 |
| 2013 | Exploiting Cyclic Prefix in Turbo FDE Systems Using Factor GraphabstractThis paper investigates the MMSE-based frequency domain equalization (FDE) algorithms in turbo equalization systems. As opposed to the conventional FDE systems where the cyclic prefix (CP) is discarded at the receiver, we take advantage of the redundancy and make use of all the observed signals for equalization purpose. First, we interpret the conventional frequency domain equalizer as a Forney-style factor graph (FFG), and accordingly an equalization algorithm is derived based on the Gaussian message passing (GMP) technique. Second, the normally discarded CP part is presented similarly using an FFG, and an algorithm that integrates both of the FFGs is proposed. As a result, two extrinsic messages about the data symbols are obtained rather than one, and they are merged together based on a symbol-wise combination. Third, with approximations made on two covariance matrices, the complexity of the proposed equalization algorithm is maintained at the same order as that of the conventional FDE algorithm, i.e. O(Nlog2N) per block per iteration. Simulations results verify that, a gain of around 0.7dB is achieved compared with the conventional algorithm at 1/4 CP ratio, for both 16QAM and 64QAM system with Gray mapping over AWGN or ISI channels. Jindan Yang, Qinghua Guo 0001, Defeng Huang, Sven Nordholm |
WCNC | 4 |
| 2013 | Speech enhancement strategy for speech recognition microcontroller under noisy environments
Kit Yan Chan, Sven Nordholm, Ka Fai Cedric Yiu, Roberto Togneri |
Neurocomputing | 2 |
| 2013 | Second-order blind signal separation with optimal step size
Hai Huyen Dam, Dedi Rimantho, Sven Nordholm |
Speech Commun. | 3 |
| 2013 | Optimization and evaluation of sigmoid function with a priori SNR estimate for real-time speech enhancement
Pei Chee Yong, Sven Nordholm, Hai Huyen Dam |
Speech Commun. | 2 |
| 2013 | Iterative Frequency Domain Equalization With Generalized Approximate Message PassingabstractAn iterative frequency domain equalization approach for coded single-carrier block transmissions over frequency selective channels is developed by using the recently proposed generalized approximate message passing (GAMP) algorithm. Compared with the low-complexity iterative frequency domain linear minimum mean square error (FD-LMMSE) equalization, the proposed approach can achieve significant performance gain with slight complexity increase. Qinghua Guo 0001, Defeng Huang, Sven Nordholm, Jiangtao Xi, Yanguang Yu |
IEEE Signal Process. Lett. | 3 |
| 2013 | New Insights Into Optimal Acoustic Feedback CancellationabstractIn this letter, we present new insights into the bias problem for acoustic feedback cancellation when a probe signal approach is used. The optimum solution of the feedback canceler is not the feedback path but the product of the feedback path and the sensitivity function and hence, the solution is biased. The novelty of this paper also consists of the derivation of the conditions for unbiased feedback cancellation when a probe signal is used as input to the canceler. An adequate delay in the forward path is necessary to reduce, or remove the bias term. The theoretical analysis is verified with simulation results. Carlos Renato C. Nakagawa, Sven Nordholm, Wei-Yong Yan |
IEEE Signal Process. Lett. | 2 |
| 2013 | Design of Steerable Spherical Broadband Beamformers With Flexible Sensor ConfigurationsabstractIn broadband beamformer applications with dynamically moving sources, it can be important to have a simple mechanism to steer the main-beam. It can also be desirable if the beampattern of the beamformer is invariant to the look direction. A number of design methods for such beamformers, based on the spherical harmonic transform, have been reported in the literature. However, these methods require the sensor positions to satisfy a certain condition which may conflict with practical considerations. This paper proposes a design method which obviates this restriction thus allowing for spherical arrays with arbitrary sensor configurations. Moreover, for a comparable level of performance and computational complexity to the existing spherical harmonic beamformers, the proposed beamformer requires fewer sensors. The trade-off is that the design of the beamformer now depends on the sensor positions. Other considerations, such as the effects of array mis-orientation and robustness, are also discussed in this paper and illustrated by design examples. Chiong-Ching Lai, Sven Nordholm, Yee Hong Leung |
IEEE Trans. Speech Audio Process. | 2 |
| 2013 | A Factor Graph Approach to Exploiting Cyclic Prefix for Equalization in OFDM SystemsabstractIn OFDM systems, cyclic prefix (CP) insertion and removal enables the use of a set of computationally efficient single-tap equalizers at the receiver. Due to the extra transmission time and energy, the CP causes a loss in both spectrum efficiency and power efficiency. On the other hand, as a repetition of part of the data, the CP brings extra information and can be exploited for detection. Therefore, instead of discarding the CP observation as in the conventional OFDM system, we utilize all the received signals in a soft-input soft-output equalizer of a turbo equalization OFDM system. First, the models for both the CP part and the non-CP part of observation are presented in a Forney-style factor graph (FFG). Then based on the computation rules of the FFG and the Gaussian message passing (GMP) technique, we develop an equalization algorithm. With proper approximation, the complexity of the proposed algorithm is reduced to \changedmathcal O(2RNlog2N+4RGlog2G+2RG) per data block for R iterations, where N is the length of the data block and G is equal to P+L-1 with P the length of the CP and L the maximum delay spread of the channel. To justify the performance improvement, SNR analysis is provided. Simulation results show that the proposed approach achieves a significant gain over the conventional approach and the turbo equalization system converges within two iterations. Jindan Yang, Qinghua Guo 0001, Defeng Huang, Sven Nordholm |
IEEE Trans. Commun. | 4 |
| 2012 | Multichannel filters for speech recognition using a particle swarm optimizationabstractSpeech recognition has been used in various real-world applications such as automotive control, electronic toys, electronic appliances etc. In many applications involved speech control functions, a commercial speech recognizer is used to identify the speech commands voiced out by the users and the recognized command is used to perform appropriate operations. However, users' commands are often corrupted by surrounding ambient noise. It decreases the effectiveness of speech recognition in order to implement the commands accurately. This paper proposes a multichannel filter to enhance noisy speech commands, in order to improve accuracy of commercial speech recognizers which work under noisy environment. An innovative particle swarm optimization (PSO) is proposed to optimize the parameters of the multichannel filter which intends to improve accuracy of the commercial speech recognizer working under noisy environment. The effectiveness of the multichannel filter was evaluated by interacting with a commercial speech recognizer, which was worked in a warehouse. Kit Yan Chan, Sven Nordholm, Ka Fai Cedric Yiu |
ICARCV | 2 |
| 2012 | A robust approach to reverberant blind source separation in the presence of noise for arbitrarily arranged sensorsabstractConsiderable attention has been devoted to the reverberant blind source separation problem: in particular, the concept of time-frequency masking. However, realistic acoustic scenarios often comprise not only reverberation, but also additive noise due to factors such as non-ideal channels. This paper presents robust evaluations of a time-frequency masking approach for separation in such realistic conditions. The fuzzy c-means clustering algorithm is used to cluster spatial feature cues into a time-frequency mask. Experimental results demonstrated superiority in separation, with notable improvements in the SNR additionally observed. Not only does this establish the proposed scheme viable for reverberant blind source separation, but also as a credible means of speech enhancement in the presence of additive broadband noise. Ingrid Jafari, Roberto Togneri, Sven Nordholm |
ICASSP | 3 |
| 2012 | Dual microphone solution for acoustic feedback cancellation for assistive listeningabstractThe method proposed in this paper improves the identification and cancellation of the feedback path by making the adaptive canceller robust against the impact of the desired speech signal. The proposed method allows for the canceller's coefficients to be continuously adapted allowing it to track variations in the feedback path even in the presence of the desired signal. It suggests the use of dual microphones and dual adaptive filters arranged in such a way that allows the speech signal to be identified and removed from the adaptation process. This results in a more robust solution which was verified by our experiments and evaluations. The perceptual evaluation of speech quality (PESQ) measure was also used to show that the proposed method results in better signal quality. Carlos Renato C. Nakagawa, Sven Nordholm, Wei-Yong Yan |
ICASSP | 2 |
| 2012 | Trade-off evaluation for speech enhancement algorithms with respect to the a priori SNR estimationabstractIn this paper, a modified a priori SNR estimator is proposed for speech enhancement. The well-known decision-directed (DD) approach is modified by matching each gain function with the noisy speech spectrum at current frame rather than the previous one. The proposed algorithm eliminates the speech transient distortion and reduces the impact from the choice of the gain function towards the level of smoothing in the SNR estimate. An objective evaluation metric is employed to measure the trade-off between musical noise, noise reduction and speech distortion. Performance is evaluated and compared between a modified sigmoid gain function, the state-of-the-art log-spectral amplitude estimator and the Wiener filter. Simulation results show that the modified DD approach performs better in terms of the trade-off evaluation. Pei Chee Yong, Sven Nordholm, Hai Huyen Dam |
ICASSP | 2 |
| 2012 | A Spectral Slit Approach to Doubletalk DetectionabstractA new class of doubletalk detector based on exploiting a spectral slit is proposed. This is achieved by spectrally deleting a frequency band in the far-end signal such that when the near-end signal is present, only the near-end spectral information is present. The proposed method relies solely on the detection of speech activity period in the slit area, and significantly, it requires no estimation of the echo path. Evaluation in typical acoustic echo setups shows that the proposed method outperforms other conventional doubletalk detectors in terms of probability of miss detection even under poor echo-to-noise ratio (ENR), low echo-to-far-end ratio (EFR) conditions, and echo path change. Siow Yong Low, Svetha Venkatesh, Sven Nordholm |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | A Fourier Based Method for Approximating the Joint Detection Probability in MIMO CommunicationsabstractWe propose a numerically efficient technique to approximate the joint detection probability of a coherent multiple input multiple output (MIMO) receiver in the presence of inter-symbol interference (ISI) and additive white Gaussian noise (AWGN). This technique approximates the probability of detection by numerically integrating the product of the characteristic function (CF) of the received filtered signal with the Fourier transform of the multi-dimension decision region. Naturally, the accuracy of the approximation is dependent on the number of points in the numerical integration. The proposed method selects the number of points to integrate over by deriving bounds on the approximation error. As such, a reduction of up to 1/100thof the number points can be achieved when comparing with existing single input single output (SISO) techniques. Simulation results demonstrate that the proposed method is highly accurate with the calculated errors being well within the bounded approximation error. In addition, the results are closely matched with those obtained using the Monte Carlo simulations with a typical error less than 10-4while requiring only a fraction of the computation time. Greg Day, Sven Nordholm, Hai Huyen Dam |
IEEE Trans. Commun. | 2 |
| 2012 | Enhancement of Speech Recognitions for Control Automation Using an Intelligent Particle Swarm OptimizationabstractFor over two decades, speech control mechanisms have been widely applied in manufacturing systems such as factory automation, warehouse automation, and industrial robotic control for over two decades. To implement speech controls, a commercial speech recognizer is used as the interface between users and the automation system. However, users' commands are often contaminated by environmental noise which degrades the performance of speech recognition for controlling automation systems. This paper presents a multichannel signal enhancement methodology to improve the performance of commercial speech recognizers. The proposed methodology aims to optimize speech recognition accuracy of a commercial speech recognizer in a noisy environment based on a beamformer, which is developed by an intelligent particle swarm optimization. It overcomes the limitation of the existing signal enhancement approaches whereby the parameters inside commercial speech recognizers are required to be tuned, which is impossible in a real-world situation. Also, it overcomes the limitation of the existing optimization algorithm including gradient descent methods, genetic algorithms and classical particle swarm optimization that are unlikely to develop optimal beamformers for maximizing speech recognition accuracy. The performance of the proposed methodology was evaluated by developing beamformers for a commercial speech recognizer, which was implemented on warehouse automation. Results indicate a significant improvement regarding speech recognition accuracy. Kit Yan Chan, Ka Fai Cedric Yiu, Tharam S. Dillon, Sven Nordholm, Sai-Ho Ling |
IEEE Trans. Ind. Informatics | 4 |
| 2011 | Performance Analysis of Enhanced Verification-Based Decoding for Packet-Based LDPC Codes over Binary Symmetric ChannelabstractIn this paper, frame error rate (FER) performance of enhanced verification-based decoding algorithm (EVA) is investigated for packet-based low-density parity-check (LDPC) codes over binary symmetric channel (BSC). By taking the false verification into account, a recursive statistical model is proposed to analyze the FER performance based on the packet-level analysis. Simulation results demonstrate that the proposed model works for various packet sizes and channel parameters of practical interest with different verification constraints. Defeng Huang, Sven Nordholm, Ying-Jun Angela Zhang |
GLOBECOM | 3 |
| 2011 | Design of robust steerable broadband beamformers incorporating microphone gain and phase error characteristicsabstractBeamformers are known to be sensitive to errors and mismatches in their array elements. This paper proposes a robust steerable broad band beamformer design using the Farrow structure and for arbitrary array geometry. The design formulation includes stochastic models describing the microphone characteristics as random variables, thus allowing flexibility for microphone errors. This method establishes a direct relationship in controlling the robustness specification in the design from given microphone characteristics. The robust design procedure optimises the mean performance of the beamformer. Design examples show significant reduction in error sensitivity in the robust design formulation. Chiong-Ching Lai, Sven Nordholm, Yee Hong Leung |
ICASSP | 2 |
| 2011 | Underdetermined Blind Source Separation with Fuzzy Clustering for Arbitrarily Arranged SensorsabstractRecently, the concept of time-frequency masking has developed as an important approach to the blind source separation problem, particularly when in the presence of reverberation. However, previous research has been limited by factors such as the sensor arrangement and/or the mask estimation technique implemented. This paper presents a novel integration of two established approaches to BSS in an effort to overcome such limitations. A multidimensional feature vector is extracted from a non-linear sensor arrangement, and the fuzzy c-means algorithm is then applied to cluster the feature vectors into representations of the source speakers. Fuzzy time-frequency masks are estimated and applied to the observations for source recovery. The evaluations on the proposed study demonstrated improved separation quality over all test conditions. This establishes the potential of multidimensional fuzzy c-means clustering for mask estimation in the context of blind source separation Ingrid Jafari, Serajul Haque, Roberto Togneri, Sven Nordholm |
INTERSPEECH | 4 |
| 2011 | A New Evidence Model for Missing Data Speech Recognition With Applications in Reverberant Multi-Source EnvironmentsabstractConventional hidden Markov model (HMM) decoders often experience severe performance degradations in practice due to their inability to cope with uncertain data in time-varying environments. In order to address this issue, we propose the bounded-Gauss-Uniform mixture probability density function (pdf) as a new class of evidence model for missing data speech recognition. Exemplary for a hands-free speech recognition scenario, we illustrate how the parameters of the new mixture pdf can be estimated with the help of a multi-channel source separation front-end. In comparison with other models the new evidence pdf retains a fuller description of the available data and provides a more effective link between source separation and recognition. The superiority of the bounded-Gauss-Uniform mixture pdf over conventional approaches is demonstrated for a connected digits recognition task under varying test conditions. Marco Kühne, Roberto Togneri, Sven Nordholm |
IEEE Trans. Speech Audio Process. | 3 |
| 2010 | A novel fuzzy clustering algorithm using observation weighting and context information for reverberant blind speech separation
Marco Kühne, Roberto Togneri, Sven Nordholm |
Signal Process. | 3 |
| 2009 | Performance Analysis of Verification-Based Decoding for Packet-Based LDPC Codes over Binary Symmetric ChannelabstractIn this paper, we propose a statistical model to analyze the performance of verification-based algorithm (VA) for packet-based low-density parity-check (LDPC) codes over binary symmetric channel (BSC). In contrast to the analysis of VA in the literature, we propose to take the false verification into consideration. For a given ensemble of LDPC codes and channel parameters, the proposed analysis model gives an efficient way to find the average performance of packet-based LDPC codes with verification-based decoding. Through numerical results, we find that the proposed method can provide a close estimation of frame error rate (FER) for packet-based LDPC codes with verification- based decoding over BSC for all crossover probabilities of practical interests. Defeng Huang, Sven Nordholm |
ICC | 3 |
| 2009 | Speech Recognition Enhancement Using Beamforming and a Genetic AlgorithmabstractThis paper proposes a genetic algorithm (GA) based beamformer to optimize speech recognition accuracy for a pretrained speech recognizer. The proposed beamformer is designed to tackle the non-differentiable and non-linear natures of speech recognition by employing the GA algorithm to search for the optimal beamformer weights. Specifically, a population of beamformer weights is reproduced by crossover and mutation until the optimal beamformer weights are obtained. Results show that the speech recognition accuracies can be greatly improved even in noisy environments. Kit Yan Chan, Siow Yong Low, Sven Nordholm, Ka Fai Cedric Yiu, Sai-Ho Ling |
NSS | 3 |
| 2009 | Robust Source Localization in Reverberant Environments Based on Weighted Fuzzy ClusteringabstractSuccessful localization of sound sources in reverberant enclosures is an important prerequisite for many spatial signal processing algorithms. We investigate the use of a weighted fuzzyc-means cluster algorithm for robust source localization using location cues extracted from a microphone array. In order to increase the algorithm's robustness against sound reflections, we incorporate observation weights to emphasize reliable cues over unreliable ones. The weights are computed from local feature statistics around sound onsets because it is known that these regions are least affected by reverberation. Experimental results illustrate the superiority of the method when compared with standard fuzzy clustering. The proposed algorithm successfully located two speech sources for a range of angular separations in room environments with reverberation times of up to 600 ms. Marco Kühne, Roberto Togneri, Sven Nordholm |
IEEE Signal Process. Lett. | 3 |
| 2008 | Enhanced Verification-Based Decoding for Packet-Based LDPC Codes over Wireless ChannelsabstractIn this paper, we apply the enhanced verification-based decoding algorithm (EVA) for packet-based LDPC codes over wireless channels. Compared with the verification algorithm (VA) in the literature, the EVA algorithm enhances the verification condition thereby reducing the likelihood of false verification. The wireless channels in simulations are modeled by binary symmetric channel (BSC) and Gilbert-Elliott (GE) channel. The numerical results indicate the proposed algorithm gives a superior performance but with only moderate computation increase compared with VA. For example, when the bit error probability is less than 10-3in bad state of GE channel, the EVA reduces the frame error rate (FER) by one order with less than 35% complexity increase compared with VA. Defeng Huang, Sven Nordholm |
GLOBECOM | 3 |
| 2008 | Towards the use of full covariance models for missing data speaker recognitionabstractThis work investigates the use of missing data techniques for noise robust speaker identification. Most previous work in this field relies on the diagonal covariance assumption in modeling speaker specific characteristics via Gaussian mixture models. This paper proposes the use of full covariance models that can capture linear correlations among feature components. This is of importance for missing data marginalization techniques as they depend on spectral rather than cepstral feature representations. Bounded and complete marginalization schemes are investigated both with diagonal and full covariance mixture models. Speaker identification experiments using stationary and non-stationary noise confirm that full covariance models are indeed superior compared to diagonal models. Marco Kühne, Daniel Pullella, Roberto Togneri, Sven Nordholm |
ICASSP | 4 |
| 2008 | Adaptive speech enhancement with varying noise backgroundsabstractWe present a new approach for speech enhancement in the presence of non-stationary and rapidly changing background noise. A distributed microphone system is used to capture the acoustic characteristics of the environment. The input of each microphone is then classified either as speech or one of the predetermined noise types. Further enhancement of speech in respective microphones is carried out using a modified spectral subtraction algorithm that incorporates multiple noise models to quickly adapt to rapid background noise changes. Tests on real world speech captured under diverse conditions demonstrate the effectiveness of this method. Thorsten Kühnapfel, Tele Tan, Svetha Venkatesh, Sven Nordholm, Burkhard Igel |
ICPR | 4 |
| 2008 | Adaptive beamforming and soft missing data decoding for robust speech recognition in reverberant environmentsabstractAbstract This paper presents a novel approach to combine microphonearray processing and robust speech recognition for reverberantmulti-speaker environments. Spatial cues are extracted from amicrophone array andautomaticallyclusteredtoestimate local-ization masks in the time-frequency domain. The localizationmasks are then used to blindly design adaptive filters in order toenhancethesourcesignalspriortomissingdataspeechrecogni-tion. A novel evidence model better exploiting the informationprovided by the source separation stage is proposed. Recogni-tion experiments demonstrate the effectiveness of the schemewhen compared to traditional microphone array enhancementand a related binaural separation model. Index Terms : missing data speech recognition, microphone ar-ray processing, adaptive beamforming 1. Introduction A key requirement for automatic speech recognition (ASR)technology to be employed in everyday situations is robust-ness to multiple competing speakers in reverberant enclosures.While today’s ASR systems still fall short in comparison withhuman listeners some promising success has been achieved indealing with multiple speakers in anechoic conditions. For ex-ample, the work reported in [1] utilizes binaural localizationcuessuchasinterauraltimeandintensitydifferencestoestimatetime-frequency (TF) masks for missing data speech recognition(MD-ASR).However,inreverberantenclosuresthelocalizationcues become increasingly unreliable making an accurate esti-mation of the TF mask considerably more difficult. Further-more when the speech models are trained on anechoic data thespectral features employed in MD-ASR are adversely affectedby the room reflections.Previous attempts to handle multisource reverberant en-vironments include dereverberation filters [2], adaptation ofspeech models to the room [3], an inhibition mechanism to em-phasizesoundonsets[4]andfeatureenhancementusinganane-choic speech prior based reconstruction technique [5].This paper proposes an alternative approach by extend-ing the two-channel system developed in [6] to the multi-microphone case (see Figure 1a). The system consists of asource separation stage coupled with a missing data decoderthrough the probabilistic concept of evidence models [7]. Forthe source separation stage a fuzzy clustering approach is em-ployed using direction of arrival (DOA) values of the two outersensors in the array. The clustering produces a source DOAestimate and a TF mask marking the dominant TF points foreach source. While in [6] the correlation between adjacent TFpoints wasonly utilized duringmask post-processing inthispa-per the neighborhood information is integrated during the clus-tering process itself. This biases the solution towards homoge-nous masks and helps greatly to reduce the effect of noise visi-ble in the localization cues under reverberant conditions.Motivated by [8] the masks and source DOAs are then uti-lized to blindly design multiple adaptive beamformers capableofenhancingthespeechsourcespriortorecognition. Thisworkproposes a novel evidence model incorporating both localiza-tion information and feature uncertainty. The system operatesinanunsupervised manneranddependsneither ontrainingdatafor source localization nor a priori learned echoic speech mod-els adapted to the reverberant environment.The remainder of this paper is as follows: Section 2 de-scribes the source separation stage in more detail. Section 3discusses the missing data recognizer and the evidence modelparameter estimation. Section 4 reports the evaluation resultsand compares these with a related binaural model. The papercloses in Section 5 with an outlook into future work. Marco Kühne, Roberto Togneri, Sven Nordholm |
INTERSPEECH | 3 |
| 2008 | Second-Order Blind Signal Separation for Convolutive Mixtures Using Conjugate GradientabstractThis letter presents a new computational procedure for the second-order gradient-based blind signal separation (BSS) problem with convolutive mixtures that has improved convergence characteristics over the steepest descent algorithm. The BSS problem is formulated as a constrained optimization problem with complex unmixing weight matrices where the constraints are formulated to overcome the permutation effects. This problem is then transformed into an unconstrained optimization problem, so that the conjugate gradient algorithm can be applied. The convergence of the proposed procedure is compared with the steepest descent algorithms in real and simulated environments. Hai Huyen Dam, Antonio Cantoni, Sven Nordholm, Kok Lay Teo |
IEEE Signal Process. Lett. | 3 |
| 2008 | Noise Statistics Update Adaptive Beamformer With PSD Estimation for Speech Extraction in Noisy EnvironmentabstractThis paper addresses the problem of extracting a desired speech source from a multispeaker environment in the presence of background noise. A new adaptive beamforming structure is proposed for this speech enhancement problem. This structure incorporates power spectral density (PSD) estimation of the speech sources together with a noise statistics update. An inactive-source detector based on minimum statistics is developed to detect the speech presence and to track the noise statistics. Performance of the proposed beamformer is investigated and compared to the minimum variance distortionless response (MVDR) beamformer with or without a postfilter in a real hands-free communication environment. Evaluations show that the proposed beamformer offers good interference and noise suppression levels while maintaining low distortion of the desired source. Hai Huyen Dam, Hai Quang Dam, Sven Nordholm |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Mel-Spectrographic Mask Estimation for Missing Data Speech Recognition using Short-Time-Fourier-Transform Ratio EstimatorsabstractThis paper adopts the framework of DUET, a recently proposed blind source separation (BSS) method, for speech recognition. Based on the attenuation and delay estimation in stereo signals spectrographic masks are designed to extract a target speaker from a mixture containing multiple speech sources. Instead of using these masks for resynthesis we avoid source reconstruction and propose to combine the source separation with a missing data speech recognizer. The obtained results for connected digit experiments in a multi-speaker environment demonstrate the validity of the approach. Marco Kühne, Roberto Togneri, Sven Nordholm |
ICASSP (4) | 3 |
| 2007 | Smooth soft mel-spectrographic masks based on blind sparse source separationabstractAbstract This paper investigates the use of DUET, a recently proposedblind source separation method, as front-end for missing dataspeech recognition. Based on the attenuation and delay estima-tion in stereo signals soft time-frequency masks are designedto extract a target speaker from a mixture containing multiplespeech sources. A postprocessing step is introduced in order toremove isolated mask points that can cause insertion errors inthe speech decoder. The results for connected digit experimentsin a multi-speaker environment demonstrate that the proposedsoft masks closely match the performance of the oracle maskdesigned with a priori knowledge of the source spectra. Index Terms : speech recognition, missing data, attenuationand delay estimation 1. Introduction The concept of time-frequency (TF) masking has recently at-tracted some interest in the field of blind signal separation (BSS)[1, 2]. Demixing via TF-masks has the potential to separatemixtures with more sources than sensors as it does not rely onmatrix inversion. Instead the TF-plane is partitioned into dis-joint regions each assigned to a particular source. The sourcesare then recovered by converting each region back into the timedomain. It seems promising to use BSS systems as front-endsfor automatic speech recognition (ASR). In [3] we have pro-posed such a combination using a BSS technique called DUETand a missing data (MD) speech recognizer. The proposedsystem uses DUET to estimate TF-masks in the sparse Short-Time-Fourier-Transform (STFT) domain before converting thehigh STFT frequency resolution to a perceptual mel-frequencyscale suitable for ASR. In this way we can avoid source re-construction and directly exploit the spectrographic masks forMD-ASR. This paper extends our previous work in two regards.Firstly, we replace binary masks with soft masks which havebeen proven to be beneficial for both speech recognition andspeech enhancement. Several studies [2, 4] have reported thatbinary TF-masking can lead to audible unnatural sound arti-facts that can be avoided to some degree by soft masks. Theadvantages of soft masks in MD-ASR are even more evidentas marginalization approaches based on soft decisions consis-tently outperformed hard masks [5]. Secondly, we show thata simple mask postprocessing can lead to substantial recogni-tion improvements. A two-dimensional (2-D) median filter wasapplied to reduce the influence of outliers visible as scattered Marco Kühne, Roberto Togneri, Sven Nordholm |
INTERSPEECH | 3 |
| 2007 | Feature and distribution normalization schemes for statistical mismatch reduction in reverberant speech recognitionabstractReverberant noise has been a major concern in speech recognition systems. Many speech recognition systems, even with state-of-art features, fail to respond to reverberant effects and the recognition rate deteriorates. This paper explores the significance of normalization strategies in reducing statistical mismatches for robust speech recognition in reverberant environment. Most normalization works focused only on ambient noise and have yet been experimented on reverberant noise. In addition, we propose a new approach for the odd order cepstral moment normalization which is computationally more efficient and reduces the convergence rate in the algorithm. The proposed method is experimentally justified and corroborated by the performance of other normalization schemes. The results emphasize the significance of reducing statistical mismatches in feature space for reverberant speech recognition. A. M. Toh, Roberto Togneri, Sven Nordholm |
INTERSPEECH | 3 |
| 2006 | Speech Signal Extraction Utilizing PCA-ICA Algorithm With a Non-Uniform Spacing Microphone ArrayabstractSpeech signal extraction is becoming more and more important as evidently displayed by its numerous applications such as mobile phones, conference equipments and surveillance. This paper presents a blind method to enhance a speech source of interest in noisy environments. The proposed technique consists of the principal component analysis (PCA) and the independent component analysis (ICA) to extract the speech signal. In an effort to overcome the small phase resolution due to the constraint on the inter-element distance, a non-uniform spacing PCA-ICA algorithm is suggested. By utilizing a different inter-element distance processing on each pair of microphones in a multistage fashion, a better separation is achieved. Results show better separation performance for the proposed method compared to the uniformly spaced microphone array. Sven Nordholm, Siow Yong Low |
ICASSP (5) | 1 |
| 2006 | A robust transform domain echo canceller employing a parallel filter structure
Jiaquan Huo, Ka Fai Cedric Yiu, Sven Nordholm, Kok Lay Teo |
Signal Process. | 3 |
| 2006 | A hybrid method for the design of oversampled uniform DFT filter banks
Ka Fai Cedric Yiu, Nedelko Grbic, Sven Nordholm, Kok Lay Teo |
Signal Process. | 3 |
| 2006 | Statistical voice activity detection using low-variance spectrum estimation and an adaptive thresholdabstractTraditionally, voice activity detection algorithms are based on any combination of general speech properties such as temporal energy variations, periodicity, and spectrum. This paper describes a novel statistical method for voice activity detection using a signal-to-noise ratio measure. The method employs a low-variance spectrum estimate and determines an optimal threshold based on the estimated noise statistics. A possible implementation is presented and evaluated over a large test set and compared to current modern standardized algorithms. The evaluations indicate promising results with the proposed scheme being comparable or favorable over the whole test set. Alan Davis, Sven Nordholm, Roberto Togneri |
IEEE Trans. Speech Audio Process. | 2 |
| 2005 | A subband space constrained beamformer incorporating voice activity detection [speech enhancement applications]abstractThis paper introduces a new subband adaptive space constrained beamforming structure for use in hands-free speech enhancement applications. The scheme incorporates a space constrained source model and voice activity information through the integration of a voice activity detector (VAD). The VAD information is used to estimate noise covariance information during non-speech periods and to optimally estimate the source power spectral density (PSD), which is used to provide a spectrally optimized constraint on the source. The proposed structure is evaluated in a real car environment, yielding results which compare well to the optimal Wiener solution where full knowledge of the source is known. Alan Davis, Siow Yong Low, Sven Nordholm, Nedelko Grbic |
ICASSP (3) | 3 |
| 2005 | Robust acoustic direction of arrival estimation using Root-SRP-PHAT, a realtime implementationabstractWideband robust direction of arrival estimation using microphone arrays is associated with computationally complex algorithms, unsuitable for implementation on low power devices. This paper improves on the computational complexity by devising an algorithm based on steered response interpolation combined with steered response power estimation employing root solving. The new algorithm is implemented in a PC-based realtime system and has been compared to Root-MUSIC and Far-Field SRP-PHAT. The comparison was performed with regards to computational load and robustness to noise and reverberation in a real room environment. The results show that the new algorithm is the most computationally efficient, while providing superior robustness compared to Root-MUSIC. Anders M. Johansson, Sven Nordholm |
ICASSP (4) | 2 |
| 2005 | A blind approach to joint noise and acoustic echo cancellationabstractThe paper introduces a new scheme which combines the popular blind signal separation (BSS) and a post-processor to suppress noise and acoustic echo jointly. The new L element structure uses the BSS as a front-end processor to extract the target signal spatially from the interference (noise and echo). Statistical measures are then employed to select the target signal dominant signal from the BSS outputs. The remaining L-1 BSS outputs (noise and echo dominant) and the existing far-end line echo are then used as the reference signals in an adaptive noise canceller (ANC) to enhance the target signal temporally. The novel structure bypasses the need for any a priori information whilst compensating the separation quality of the BSS temporally. Real room evaluations demonstrate the efficacy of the scheme in both noisy double-talk and non double-talk situations. Siow Yong Low, Sven Nordholm |
ICASSP (3) | 2 |
| 2005 | Frequency Domain Blind Equalization for MIMO SystemsabstractThis paper presents a new normalized frequency domain approach for adaptive blind equalization for multiple-input multiple-output (MIMO) communication systems. We first develop a time domain block based algorithm that updates the equalizer coefficients once per block of data symbols. As the time domain algorithm involves infinite summations in the separation cost function, an approximation is then proposed to reduce the cost function to finite summations. As a consequence, the block based time domain algorithm can be implemented in the frequency domain to reduce the computational complexity associated with the conventional symbol-by-symbol time domain update. Furthermore, by recognizing that the signals in the frequency domain are orthogonal, it is possible to significantly improve the convergence rate by normalizing the update equations with respect to the signal power in each frequency bin. Simulation results show that the proposed algorithm can successfully equalize the received signals while maintaining low computational complexity. Moreover, the presented normalized frequency domain blind equalization algorithm for MIMO systems significantly improves the convergence rate. Hai Huyen Dam, Sven Nordholm, Hans-Jürgen Zepernick |
PIMRC | 2 |
| 2005 | Optimized decision delays in finite-length MIMO DFEabstractFinite-length multiple-input multiple-output (MIMO) decision feedback equalizers (DFEs) that optimize the decision delays are investigated herein. The mean square error (MSE) of each output with respect to the decision delays is derived to form the basis of the optimization cost function. A genetic algorithm (GA) is implemented to reduce the computational burden of the search. Simulation results show that the GA successfully tracks the optimum decision delays with only a fraction of the computational complexity of the global search. Moreover, the use of optimized decision delays in the MIMO DFE for simulated fading MIMO channels provides an average improvement of 4.1 dB for the combined output MSE. Hai Huyen Dam, Sven Nordholm, Greg Day |
IEEE Signal Process. Lett. | 2 |
| 2004 | Subband adaptive equalisation with optimised alignment delaysabstractThe multi-rate subband equaliser converts the computationally complex fullband equaliser structure into several parallel low complexity subband equalisers, each operating independently. The fullband equaliser has a minimum mean square error (MMSE) dependent upon the filter coefficients and the relative delay between the desired symbol and received signal. The subband equalisers MMSE performance is also dependent upon the subband adaptive filter (SAF) coefficients and the relative delay between the desired and received subband signal components, referred herein as the alignment delay. We propose a new structure allowing implementation of independent alignment delays for each subband to achieve the MMSE with respect to both the SAF coefficients and the alignment delay. Simulation results show that the use of independent alignment delays in each subband provides significant improvement over the standard subband transversal equaliser. Greg Day, Sven Nordholm, Hai Huyen Dam |
GLOBECOM | 2 |
| 2004 | Space constrained beamforming with source PSD updatesabstractThis paper presents a new space constrained adaptive beamformer employing an updated source power spectral density (PSD). The space constraints are used to capture the target signal spatially and to provide robustness against steering error vectors. The PSD update on the other hand ensures that the spectral information of the desired source is reflected continuously on the space constraints. As such, target signal extraction can be achieved with minimum distortion. The beamformer operates in a subband structure to allow time-frequency operation for each channel, yielding a combination of weighted spatial and temporal filters. Evaluations on real car data show that the proposed algorithm significantly improves the speech intelligibility with noise suppression level up to 21 dB. Hai Quang Dam, Siow Yong Low, Hai Huyen Dam, Sven Nordholm |
ICASSP (4) | 4 |
| 2004 | Spatio-temporal processing for distant speech recognitionabstractA new subband based front-end processor for speech recognition is presented. It integrates both spatial and temporal signal processing methods to enhance noisy signals as a means to reduce the mismatch problem in speech recognition. The approach makes use of the popular blind signal separation (BSS) to spatially separate the target signal from the interference. Due to the multipath/reverberant environment, BSS has its fundamental limitation in the separation quality. To overcome that, an adaptive noise canceller (ANC) is employed to perform further interference reduction. Experimental results show that even in an adverse environment, the proposed structure improves the word recognition rate (WRR) by 70% for the connected digit recognition task. Siow Yong Low, Roberto Togneri, Sven Nordholm |
ICASSP (1) | 3 |
| 2004 | Multicriteria design of oversampled uniform DFT filter banksabstractSubband adaptive filters have been proposed to avoid the drawbacks of slow convergence and high computational complexity associated with time domain adaptive filters. However, subband processing causes signal degradations due to aliasing effects and amplitude distortions. This problem is unavoidable due to further filtering operations in subbands. In this letter, the problems of aliasing effect and amplitude distortion are studied. Prototype filters which are optimized with respect to those properties are designed and their performances are compared. Moreover, the effect of the number of subbands, the oversampling factors and the length of the prototype filter are also studied. Using the multicriteria formulation, all Pareto optimums are sought via the nonlinear programming technique. We find that the prototype filter designed via the Kaiser window provides the best overall performance among the methods we studied. Also, there is a critical oversampling factor beyond which the improvement of performance is diminishing. Finally, if the length of the prototype filter increases with the number of subbands, an increase in the number of subbands will not deteriorate the performance. Ka Fai Cedric Yiu, Nedelko Grbic, Sven Nordholm, Kok Lay Teo |
IEEE Signal Process. Lett. | 3 |
| 2004 | Convolutive blind signal separation with post-processingabstractA new subband based speech enhancement scheme is presented. It integrates spatial and temporal signal processing methods to enhance speech signals in a noisy environment. The approach makes use of the popular blind signal separation (BSS) to spatially separate the target signal from the interference. Due to the multipath/reverberant environment, BSS has its fundamental limitation in its separation quality. To overcome that, an adaptive noise canceller (ANC) is employed to perform further interference reduction. The reference for the ANC in this case is simply the interference dominant output from the BSS. A higher order statistical method is proposed for the selection of the reference signal. This post processing acts as a spectral decorrelator and experimental results show that even in under-determined (more sources than elements) case, the structure offers impressive enhancement capability. Further, a remarkable improvement in recognition rate is registered when tested in automatic speech recognition (ASR). Siow Yong Low, Sven Nordholm, Roberto Togneri |
IEEE Trans. Speech Audio Process. | 2 |
| 2003 | Frequency domain constant modulus algorithm for broadband wireless systemsabstractBroadband wireless systems require powerful equalizers which can cope with long channel impulse responses and severe intersymbol interference (ISI). In this paper, a frequency domain modified constant modulus algorithm (PD-MCMA) is presented. The employed modified CMA (MCMA) is chosen to confine the phase ambiguity of the conventional CMA to a 90/spl deg/ phase shift solution. On this foundation, a block MCMA algorithm which updates the equalizer coefficients on a block-by-block basis is proposed. The advantage of this block MCMA is that it can be equivalently implemented in the frequency domain. This approach offers a large reduction in computational complexity compared to time domain algorithms. For an equalizer with 256 coefficients, the number of real multiplications required by the FD-MCMA is only 9.5% of that required by the MCMA. Furthermore, convergence of the proposed FD-MCMA is significantly improved by properly normalizing the update equations. Simulations show that the normalized FD-MCMA converges much faster than the MCMA and the FD-MCMA while maintaining approximately the same bit error rate (BER) performance. Hai Huyen Dam, Sven Nordholm, Hans-Jürgen Zepernick |
GLOBECOM | 2 |
| 2003 | Speech enhancement using multiple soft constrained subband beamformers and non-coherent techniqueabstractThis paper presents a new robust microphone array processing technique to enhance speech signals under the influence of noise and jammer(s). The new structure comprises two soft constrained subband beamformers and a non-coherent processing technique. Essentially, the first beamformer enhances the desired speech signal in a specified constrained region. The residual interference in the beamformer's output is then spectral subtracted using the estimated interference from the second beamformer. Evaluations in a real office environment show higher interference suppression compared to those obtained using the soft constrained beamformer only. Most importantly, this is achieved with negligible expense on target signal distortion. Siow Yong Low, Nedelko Grbic, Sven Nordholm |
ICASSP (5) | 3 |
| 2003 | Optimal FIR subband beamforming for speech enhancement in multipath environmentsabstractThe letter provides an analysis of optimal finite-impulse response subband beamforming for speech enhancement in multipath environments. A modification of the direct-path standard Wiener formulation is shown to give nearly optimal performance when it comes to signal-to-noise-plus-interference ratio, by including coherent multipath propagation in the criterion. Nedelko Grbic, Sven Nordholm, Antonio Cantoni |
IEEE Signal Process. Lett. | 2 |
| 2003 | Filter bank design for subband adaptive microphone arraysabstractThis paper presents a new method for the design of oversampled uniform DFT-filter banks for the special application of subband adaptive beamforming with microphone arrays. Since array applications rely on the fact that different source positions give rise to different signal delays, a beamformer alters the phase information of the signals. This in turn leads to signal degradations when perfect reconstruction filter banks are used for the subband decomposition and reconstruction. The objective of the filter bank design is to minimize the magnitude of all aliasing components individually, such that aliasing distortion is minimized although phase alterations occur in the subbands. The proposed method is evaluated in a car hands-free mobile telephony environment and the results show that the proposed method offers better performance regarding suppression levels of disturbing signals and much less distortion to the source speech. Jan Mark de Haan, Nedelko Grbic, Ingvar Claesson, Sven Nordholm |
IEEE Trans. Speech Audio Process. | 4 |
| 2003 | Performance limits in subband beamformingabstractThis paper analyzes subband beamforming schemes mainly aimed at speech enhancement and acoustic echo suppression applications such as hands-free telephony for both mobile and office environments, Internet telephony and video conferencing. Analytical descriptions of both causal finite-length and noncausal infinite-length subband microphone array structures are given. More specifically, this paper compares finite Wiener filter performance with the noncausal Wiener solution, giving a comprehensive theoretical suppression limit. It is shown that even short filters will yield a good approximation of the infinite solution, provided that the element spacing and temporal sampling is matched to the frequency band of interest. Typically, 10-20 FIR taps are sufficient in each subband. Sven Nordholm, Ingvar Claesson, Nedelko Grbic |
IEEE Trans. Speech Audio Process. | 1 |
| 2003 | Near-field broadband beamformer design via multidimensional semi-infinite-linear programming techniquesabstractBroadband microphone arrays has important applications such as hands-free mobile telephony, voice interface to personal computers and video conference equipment. This problem can be tackled in different ways. In this paper, a general broadband beamformer design problem is considered. The problem is posed as a Chebyshev minimax problem. Using the l/sub 1/-norm measure or the real rotation theorem, we show that it can be converted into a semi-infinite linear programming problem. A numerical scheme using a set of adaptive grids is applied. The scheme is proven to be convergent when a certain grid refinement is used. The method can be applied to the design of multidimensional digital finite-impulse response (FIR) filters with arbitrarily specified amplitude and phase. Ka Fai Cedric Yiu, Xiaoqi Yang 0001, Sven Nordholm, Kok Lay Teo |
IEEE Trans. Speech Audio Process. | 3 |
| 2002 | Design of spreading sequences using a modified bridging methodabstractThe performance of a code division multiple access system depends on the correlation properties of the employed spreading code. An auto-correlation function with a distinct peak enables proper synchronization and suppresses intersymbol interference. Low cross-correlation values between spreading sequences are desired to suppress multiple access interference. However, these requirements contradict each other and a trade-off needs to be established. In this paper, a modified bridging method is proposed to minimize cross-correlation with auto-correlation being allowed to lie within a fixed region. This approach is applied to the design of complex spreading sequences. Hai Huyen Dam, Hans-Jürgen Zepernick, Sven Nordholm, Jörgen Nordberg |
GLOBECOM | 3 |
| 2002 | Non-causal delayless subband adaptive equalizerabstractIn wireless communication, multiple versions of the transmitted signals arrive at the receiver with different attenuations and time delays. Thus, an equalizer with long filter is required at the receiver to reverse the multipath effects. In this paper;, a new non-causal delayless subband equalizer structure is proposed to reduce the computational complexity of the equalizer and to improve the channel tracking capability. This structure avoids the additional delays usually associated with the subband schemes. A design method is formulated to minimize the aliasing effect of the filter bank in the subbands. The significance of employing a non-causal equalizer is also emphasized. Simulation results show that the performance of the new structure is significantly improved by using the design method. Moreover, the complexity of the new subband equalizer is a fraction of the equivalent fullband at the expense of small degradation in the performance. Hai Huyen Dam, Jörgen Nordberg, Sven Nordholm |
ICASSP | 3 |
| 2002 | Soft constrained subband beamforming for hands-free speech enhancementabstractThis paper introduces a new constrained adaptive subband beamformer algorithm for speech enhancement in acoustic telecommunication systems. The solution relies on a pre-calculated source covariance matrix and recursive estimates of background noise- and handsfree signal covariance matrices. The constraint acts as an eye-opening in a vicinity of the near-field location of the source and degradations from steering-vector errors can therefor be made small. The algorithm is applied in subbands using a uniform multi channel over-sampled filterbank. Simulations with real speech recorded in an automobile hands-free environment show 19 dB noise reduction and 20 dB hands-free suppression. Nedelko Grbic, Sven Nordholm |
ICASSP | 2 |
| 2002 | Design and evaluation of nonuniform DFT filter banks in subband microphone arraysabstractThis paper presents a method for the design of nonuniform DFT filter banks for subband beamforming. Filter banks designed with the method are evaluated in subband beamforming in a real-world microphone array application. Different source positions in array applications give rise to different signal delays, which means that adaptive beamformers in the subbands alter the phase information of the subband signals in order to extract the source from a noisy background. Phase alterations in the subbands lead to signal degradations when perfect reconstruction filter banks are used for the subband decomposition and reconstruction. The objective of the proposed design is to minimize the magnitude of all aliasing components individually, such that aliasing distortion is minimized although phase alterations occur in the subbands. The proposed method is evaluated in a car hands-free mobile telephony environment with real speech signals and the results show that the performance can be increased by several decibel when using nonuniform filter banks instead of uniform filterbanks while maintaining the length of the subband filters. Jan Mark de Haan, Nedelko Grbic, Ingvar Claesson, Sven Nordholm |
ICASSP | 4 |
| 2002 | A new design method for broadband microphone arrays for speech input in automobilesabstractA new design method for broadband microphone arrays is presented. Using sequences of calibration signals, the method is able to design finite-impulse response (FIR) filters with specific performance. The method can control and adjust the speech distortion, noise suppression, and echo cancellation directly. It turns out that a significantly shorter filter length can be applied to achieve better overall performance than the least-squares method or the signal-to-noise plus interference method. Ka Fai Cedric Yiu, Nedelko Grbic, Kok Lay Teo, Sven Nordholm |
IEEE Signal Process. Lett. | 4 |
| 2001 | New weight transform schemes for delayless subband adaptive filteringabstractVery long adaptive filters are used for controlling the acoustic echo due to the acoustic echo of the loudspeaker and microphone. Subband adaptive filters have been proposed to save computations as well as speed up convergence. Delayless subband adaptive filters remove the additional signal path delay due to the analysis-synthesis system in their conventional counterpart by transforming the subband filters to a corresponding fullband one and cancelling the echo using the fullband filter. The weight transform from subband to fullband plays a crucial role in the delayless subband adaptive filters. In this paper, a defect of the FFT-stacking weight transform is identified, and new transform schemes are proposed. Simulation results show that significant performance improvements can be achieved by employing the proposed weight transform schemes. Jiaquan Huo, Sven Nordholm, Zhuquan Zang |
GLOBECOM | 2 |
| 2001 | Design of oversampled uniform DFT filter banks with delay specification using quadratic optimizationabstractSubband adaptive filters have been proposed to avoid the drawbacks of slow convergence and high computational complexity associated with time domain adaptive filters. Subband processing introduces transmission delays caused by the filter bank and signal degradations due to aliasing effects. One efficient way to reduce the aliasing effects is to allow a higher sample rate than critically needed in the subbands and thus reduce subband signal degradation. We suggest a design method, for uniform DFT filter banks with any oversampling factor, where the total filter bank group delay may be specified, and where the aliasing and magnitude/phase distortions are minimized. Jan Mark de Haan, Nedelko Grbic, Ingvar Claesson, Sven Nordholm |
ICASSP | 4 |
| 2001 | Blind signal separation using overcomplete subband representationabstractThis paper discusses a multirate filterbank-based extended infomax algorithm for real-world signal separation, i.e., convolved mixtures separation. Since convolution in the time domain corresponds to instantaneous mixing in the frequency domain, polyphase subband projection naturally becomes an efficient alternative to the Fourier transform based frequency domain approach. The online implementation proposed is featured by a simultaneous inverse channel identification in the frequency domain and signal filtering in the time domain. It is shown that an over-representation structure reduces aliasing between different bands and results in more accurate inverse channel estimates. Therefore, it provides better performance than the Fourier transform based structure in the measures of both separation and distortion. The performance limitation of the method is also evaluated in terms of the Wiener solution. Nedelko Grbic, Xiao-Jiao Tao, Sven Nordholm, Ingvar Claesson |
IEEE Trans. Speech Audio Process. | 3 |
| 2001 | Spectral subtraction using reduced delay convolution and adaptive averagingabstractIn hands-free speech communication, the signal-to-noise ratio (SNR) is often poor, which makes it difficult to have a relaxed conversation. By using noise suppression, the conversation quality can be improved. This paper describes a noise suppression algorithm based on spectral subtraction. The method employs a noise and speech-dependent gain function for each frequency component. Proper measures have been taken to obtain a corresponding causal filter and also to ensure that the circular convolution originating from fast Fourier transform (FFT) filtering yields a truly linear filtering. A novel method that uses spectrum-dependent adaptive averaging to decrease the variance of the gain function is also presented. The results show a 10-dB background noise reduction for all input SNR situations tested in the range -6 to 16 dB, as well as improvement in speech quality and reduction of noise artifacts as compared with conventional spectral subtraction methods. Harald Gustafsson, Sven Nordholm, Ingvar Claesson |
IEEE Trans. Speech Audio Process. | 2 |
| 1999 | Spectral subtraction with adaptive averaging of the gain function
Harald Gustafsson, Sven Nordholm, Ingvar Claesson |
EUROSPEECH | 2 |
| 1999 | Adaptive microphone array employing calibration signals: an analytical evaluationabstractThis paper gives an analytical description of an adaptive microphone array that facilitates a simple built-in calibration to the environment and instrumentation. This method, suggested for use in hands-free mobile telephones and speech recognition systems for cars, provides speech enhancement and acoustic echo-cancellation. The scheme offers several advantages, such as a simple calibration procedure, suppression of directional sources, versatile robust beamforming, and reduced target signal distortion. The analysis employs noncausal Wiener filters yielding compact and effective theoretical suppression limits. Sven Nordholm, Ingvar Claesson, Mattias Dahl |
IEEE Trans. Speech Audio Process. | 1 |
| 1997 | Analytical evaluation of a self-calibrating microphone arrayabstractThis paper gives an analytical description of an adaptive microphone array which facilitates a simple built-in calibration to the environment and instrumentation. The scheme offers several advantages, such as a simple calibration procedure and reduced target signal distortion. The analysis employs noncausal Wiener filters yielding compact and effective theoretical suppression limits. Sven Nordholm, Ingvar Claesson |
ICASSP | 1 |
| 1994 | Weighted Chebyshev approximation for the design of broadband beamformers using quadratic programmingabstractA method to solve a general broadband beamformer design problem is formulated as a quadratic program. As a special case, the minimax near-field design problem of a broadband beamformer is solved as a quadratic programming formulation of the weighted Chebyshev approximation problem. The method can also be applied to the design of multidimensional digital FIR filters with an arbitrarily specified amplitude and phase. For linear phase multidimensional digital FIR filters, the quadratic program becomes a linear program. Examples are given that demonstrate the minimax near-field behavior of the beamformers designed.> Sven Nordebo, Ingvar Claesson, Sven Nordholm |
IEEE Signal Process. Lett. | 3 |